How Amaleh built this site

Every claim the landing page makes about the skill is backed by a measured figure from the two runs that built this website. Nothing here is a benchmark. These are local estimates from the records that exist.

The project's name is عمله, transliterated as ʿamalah and pronounced Ah-mah-leh. It is a Persian word meaning workers / laborers — people contributing effort to a shared result. Amaleh is the project's Latin-script name.

The two runs

Two runs built and shipped this website: website-visuals (6 tasks) and website-polish (11 tasks). The table sums both.

RunTasksModel callsReviewsBlocking findingsCheck runsCost estimate
website-visuals6321811611.04 USD
website-polish1112456342096.78 USD
Both runs171567445270~7.82 USD

Cost is pi's local estimate of the worker and reviewer calls, not a billed figure; the coordinating host model is not in it. A blocking finding is a defect an independent reviewer from a different model family refused to accept.

The skill's own output

These are real terminal screenshots of the skill reporting on the website-polish run. The first shows the readable status report. The second shows delegation health, model speed per role and estimated spend by model.

status for website-polishterminal output
Terminal output showing the status report for the website-polish run: tasks, statuses, check receipts and completion summary.
Status report for website-polish: tasks, their statuses, registered check receipts and the completion summary.
health for website-polishterminal output
Terminal output showing delegation health for website-polish: worker versus reviewer versus manual dispatch counts, model speed per role, and estimated spend per model.
Delegation health for website-polish: worker versus reviewer versus manual dispatch counts, model speed per role, and estimated spend by model.

Claim and proof

Each claim the landing page makes about Amaleh, paired with the measured figure that backs it.

Cheap Flash models do the actual work.

156 calls

156 model calls across two runs, most from DeepSeek, GLM, Solar and MiMo Flash models. Workers cost an estimated 4.47 US dollars over 57 calls in website-polish alone.

Another model family reviews every chunk.

74 reviews

74 independent reviews, each from a different model family than the author. The reviewer never sees the author's proposed verdict.

Defects are caught before merge.

45 blocking

45 blocking review findings, every one repaired inside the chunk before it was accepted. A sixth of the total reviews found something worth blocking.

The run resumes from disk.

270 checks

270 check runs across both runs. Durable checkpoints mean an interrupted session picks up exactly where it stopped, with no work lost.

It costs little.

~7.82 USD

Estimated worker and reviewer spend: 7.82 US dollars for two full runs that built, reviewed and shipped a seven-route website. Cost is pi's local estimate, not a billed figure, and the coordinating host model is not in it.

Defects independent review caught before merge

Every one of these was blocking: a reviewer from a different model family refused to accept the chunk until the fix was in. All six were repaired inside the chunk.

before

Hero with reduced motion showed a permanently blank canvas for visitors who prefer stillness, because the accessible fallback was hidden while the canvas drew nothing.

after

The fallback now stays visible alongside the canvas. A reviewer from a different model family caught this before the chunk merged.

before

The coordinator, worker and reviewer nodes vanished almost immediately in the live animation in every WebGL browser, leaving the hero empty.

after

Node timing was fixed so the workflow it exists to show stays visible throughout. A reviewer flagged this as a blocking finding.

before

At desktop widths the "On this page" list appeared twice on every documentation page, and at 320, 390 and 768 CSS pixels it was missing entirely.

after

The list now renders exactly once at every width, in its own right rail from xl and inline below. The static checker was not catching this; a browser review did.

before

Three source files carried the wrong pronunciation and the static checker asserted the wrong value, so fixing the sources would have broken the checker.

after

Both the sources and the checker were corrected in the same chunk. A reviewer found the mismatch between what the page said and what the checker checked.

before

At 320 CSS pixels the Credentials chip overran its card and the whole landing page scrolled sideways.

after

The chip wraps properly at phone width. The layout probe at 320 pixels now passes.

before

The new two-column limit for prose cards was written into the design system but the checker never ran it, so continuous integration would have passed a page that broke the rule.

after

The prose-grid-columns gate was added to the static checker. The rule now blocks any prose card grid declaring more than two columns.

In-task decisions

Workers consulted Jev directly for bounded either-or questions without escalating to the coordinator. 27 times across both runs.

  • JevWorkers asked Jev for an in-task decision 27 times without the coordinator: 3 in website-visuals, 24 in website-polish.
  • JevJev selects proceed, research, focused grilling or brainstorming from an actual gap. No coordinator needed.
  • JevModels that did the work: DeepSeek, GLM, Solar, MiMo. Kimi handled escalated repairs.

What the runtime learned

The same run records that show the numbers also show where the cheap-model loop needed help. These are not failures; they are the reasons the runtime now refuses to finish a run while any warning stands.

Host exceptions

The coordinator took over implementation where the cheap-model loop could not finish: 9 host exceptions in website-polish and 1 in website-visuals. Each is recorded with its reason.

Health warnings

website-polish closed with 3 delegation-health warnings acknowledged: the coordinator made 35 worker, reviewer, repair and accept calls by hand; 23 invalidations reopened delivered chunks because their own checks could not see the defect; and Kimi took 35% of the estimated spend.

Kimi share

Kimi, the deeper repair model, took the largest single share of website-polish spend: an estimated 2.38 of 6.78 US dollars. That is a cost of escalation, not the baseline.

How the skill is tested

The skill is tested two ways. A live delegated fixture runs a billing scenario through the real OpenRouter credential: Jev routing, Flash implementation, an independent other-family read-only review, exact amount tests, and acceptance, integration, completion. No separate TypeSafe key is required.

An offline behavioral suite runs without any model. It covers routing, lifecycle recovery, parallel ownership, review coverage, context isolation and startup preflight. The latest development check: 83 Bun tests and 95 Node tests passed.

Skill runtime sourceVerified version
Bun1.4.2
Node.js latest stable26.9.0
pi coding agent0.85.1
Jev OpenRouter modeltypesafe/jev-1.13

The skill has no npm runtime dependency and no external skill dependencies. The runtime source executes without installing its development dependencies, the invoking agent drives the next-action loop, and the installer links the canonical skill directory into Claude and Codex.