Cheap Flash models do the actual work.
156 calls156 model calls across two runs, most from DeepSeek, GLM, Solar and MiMo Flash models. Workers cost an estimated 4.47 US dollars over 57 calls in website-polish alone.
Every claim the landing page makes about the skill is backed by a measured figure from the two runs that built this website. Nothing here is a benchmark. These are local estimates from the records that exist.
The project's name is عمله, transliterated as ʿamalah and pronounced Ah-mah-leh. It is a Persian word meaning workers / laborers — people contributing effort to a shared result. Amaleh is the project's Latin-script name.
Two runs built and shipped this website: website-visuals (6 tasks) and website-polish (11 tasks). The table sums both.
| Run | Tasks | Model calls | Reviews | Blocking findings | Check runs | Cost estimate |
|---|---|---|---|---|---|---|
| website-visuals | 6 | 32 | 18 | 11 | 61 | 1.04 USD |
| website-polish | 11 | 124 | 56 | 34 | 209 | 6.78 USD |
| Both runs | 17 | 156 | 74 | 45 | 270 | ~7.82 USD |
Cost is pi's local estimate of the worker and reviewer calls, not a billed figure; the coordinating host model is not in it. A blocking finding is a defect an independent reviewer from a different model family refused to accept.
These are real terminal screenshots of the skill reporting on the website-polish run. The first shows the readable status report. The second shows delegation health, model speed per role and estimated spend by model.


Each claim the landing page makes about Amaleh, paired with the measured figure that backs it.
Cheap Flash models do the actual work.
156 calls156 model calls across two runs, most from DeepSeek, GLM, Solar and MiMo Flash models. Workers cost an estimated 4.47 US dollars over 57 calls in website-polish alone.
Another model family reviews every chunk.
74 reviews74 independent reviews, each from a different model family than the author. The reviewer never sees the author's proposed verdict.
Defects are caught before merge.
45 blocking45 blocking review findings, every one repaired inside the chunk before it was accepted. A sixth of the total reviews found something worth blocking.
The run resumes from disk.
270 checks270 check runs across both runs. Durable checkpoints mean an interrupted session picks up exactly where it stopped, with no work lost.
It costs little.
~7.82 USDEstimated worker and reviewer spend: 7.82 US dollars for two full runs that built, reviewed and shipped a seven-route website. Cost is pi's local estimate, not a billed figure, and the coordinating host model is not in it.
Every one of these was blocking: a reviewer from a different model family refused to accept the chunk until the fix was in. All six were repaired inside the chunk.
Hero with reduced motion showed a permanently blank canvas for visitors who prefer stillness, because the accessible fallback was hidden while the canvas drew nothing.
The fallback now stays visible alongside the canvas. A reviewer from a different model family caught this before the chunk merged.
The coordinator, worker and reviewer nodes vanished almost immediately in the live animation in every WebGL browser, leaving the hero empty.
Node timing was fixed so the workflow it exists to show stays visible throughout. A reviewer flagged this as a blocking finding.
At desktop widths the "On this page" list appeared twice on every documentation page, and at 320, 390 and 768 CSS pixels it was missing entirely.
The list now renders exactly once at every width, in its own right rail from xl and inline below. The static checker was not catching this; a browser review did.
Three source files carried the wrong pronunciation and the static checker asserted the wrong value, so fixing the sources would have broken the checker.
Both the sources and the checker were corrected in the same chunk. A reviewer found the mismatch between what the page said and what the checker checked.
At 320 CSS pixels the Credentials chip overran its card and the whole landing page scrolled sideways.
The chip wraps properly at phone width. The layout probe at 320 pixels now passes.
The new two-column limit for prose cards was written into the design system but the checker never ran it, so continuous integration would have passed a page that broke the rule.
The prose-grid-columns gate was added to the static checker. The rule now blocks any prose card grid declaring more than two columns.
Workers consulted Jev directly for bounded either-or questions without escalating to the coordinator. 27 times across both runs.
The same run records that show the numbers also show where the cheap-model loop needed help. These are not failures; they are the reasons the runtime now refuses to finish a run while any warning stands.
The coordinator took over implementation where the cheap-model loop could not finish: 9 host exceptions in website-polish and 1 in website-visuals. Each is recorded with its reason.
website-polish closed with 3 delegation-health warnings acknowledged: the coordinator made 35 worker, reviewer, repair and accept calls by hand; 23 invalidations reopened delivered chunks because their own checks could not see the defect; and Kimi took 35% of the estimated spend.
Kimi, the deeper repair model, took the largest single share of website-polish spend: an estimated 2.38 of 6.78 US dollars. That is a cost of escalation, not the baseline.
The skill is tested two ways. A live delegated fixture runs a billing scenario through the real OpenRouter credential: Jev routing, Flash implementation, an independent other-family read-only review, exact amount tests, and acceptance, integration, completion. No separate TypeSafe key is required.
An offline behavioral suite runs without any model. It covers routing, lifecycle recovery, parallel ownership, review coverage, context isolation and startup preflight. The latest development check: 83 Bun tests and 95 Node tests passed.
| Skill runtime source | Verified version |
|---|---|
| Bun | 1.4.2 |
| Node.js latest stable | 26.9.0 |
| pi coding agent | 0.85.1 |
| Jev OpenRouter model | typesafe/jev-1.13 |
The skill has no npm runtime dependency and no external skill dependencies. The runtime source executes without installing its development dependencies, the invoking agent drives the next-action loop, and the installer links the canonical skill directory into Claude and Codex.