Review and recovery

What keeps a delegated run honest: a reviewer from a different model family with structured coverage, bounded repair escalation, and state that survives a cleared chat, an exhausted provider or a compacted host.

workerrouted family · writeschecksregistered · receiptsreviewerother model familyacceptall coveredflash ×2two repair cycleskimi ×1deeper diagnosishostfinal · checks onlyresumeinvoke again · same run
The review and repair loop: the worker's checks and an independent reviewer from a different model family gate acceptance; blocking findings escalate repair through flash, then Kimi, then the host, whose fix is final and accepted on its checks alone, and the checkpointed run resumes from disk on the next invocation.

Independent cross-family review

references/review.md

Review is fresh and cross-family by construction: DeepSeek writes and GLM reviews, GLM writes and DeepSeek reviews. A different provider hosting the same model family is not independence, and review candidates always exclude the author’s family; Kimi work gets a capable different family. Host work gets no model review, because a reviewer that rejects it can only send it back to the host.

The reviewer receives fresh read-only context — intent, criteria, standards, the actual changed files, relevant callers and check receipts — and never the author’s reasoning or proposed verdict. It may trace affected consumers beyond the diff.

One reviewer groups the baseline lenses: Spec/requirements, Standards/simplicity, correctness and omissions, with Spec and Standards coverage kept separate in the report. Jev selects additional lenses grounded in real changed behavior — money/error paths, contracts, security/authorization, migrations, visual/accessibility, resources, dependencies, test quality or observed performance. Sensitive behavior can get additional focused review in parallel.

Every finding needs a reachable scenario and evidence. Observations are distinguished from inference, and new defects from unrelated inherited problems. A taste preference is not a defect, and a concurrency claim needs an actual overlap, retry or shared-state path. Duplicate claims are archived and grouped without suppressing valid evidence.

Coverage before verdict

references/review.mdreferences/runtime.md

Before the first review, the host maps the actual change to requirements, callers and consumers, and reachable transitions — for an orchestrator, every operation kind from start through interrupted ownership and resume. For a simple pure calculation, an inapplicable lens is declared with a reason instead of inventing safeguards.

review-packet returns the factual contract and the exact required coverage IDs. Every review supplies coverage: exactly one entry of id, status and evidence per obligation in the packet.

StatusMeaning
coveredThe obligation was inspected and supported by named evidence.
findingA defect was identified, with a reachable scenario and evidence.
unreviewedThe obligation was not examined — this blocks acceptance.
not-applicableDeclared only with an explicit reason. Criteria and outcomes can never be not-applicable.

The runtime rejects malformed coverage and blocks acceptance while entries are missing, unreviewed or findings; old receipts without coverage remain readable but cannot authorize a new acceptance. next returns review-evidence-needed for evidence gaps, distinct from triage-repair for real blocking defects — and missing evidence does not consume repair cycles.

For a missing executable probe, the host registers one with review-check while verification is idle, runs it, and the invalidated review is replaced by a fresh corrected receipt that considers the new evidence. Read-only reviewers request probes; they do not execute commands.

Repair escalation

references/review.md

repair increments a persistent counter and chooses Flash, then Kimi, then host takeover. There is no reset through rename, a new commit or a restart, and replanning retains ancestry. The initial review is not counted as a repair cycle.

StageAllowance
Flash repairTwo ordinary repair cycles by routed Flash workers.
Kimi repairOne deeper diagnosis-and-repair cycle once the Flash allowance is exhausted.
Host takeoverThe actual invoking model diagnoses and repairs. This is the last step: no model reviews host work, and delegate accepts it once its checks pass.

Contract problems spend no repair cycles. When a worker changes files outside its task’s resources, delegate returns a scope-question instead of starting a repair. Widening the resources with amend keeps the worker’s output and verifies it again with no new worker run; delegating again reverts the paths through a repair, which does spend a cycle. An amend spends a cycle only when a check other than scope is failing or a blocking finding is open.

A worker that cannot deliver the contract as written, because a check cannot pass for a reason outside the task, two criteria contradict each other, or the work needs vendor code no source supplies, raises a contract question instead of bending the work to fit. delegate returns contract-question before any check runs, without spending a cycle. Amending the contract spends none either, and a change to checks or resources alone keeps the output for verification; delegating again with a brief that says why the contract stands sends the output to checks and review as it stands.

A reopen charges only the task it names. Its dependents carry no defect of their own: an accepted dependent keeps its output and, once its upstream is integrated again and merged into its checkout, is checked and reviewed with no worker run. A dependent whose own last attempt had already failed returns for a repair and spends a cycle.

Host work is final. What the invoking model writes under host-exception is accepted once its checks and scope pass, with no model review; a failing check sends it back to the host as host-checks-failed without spending a cycle. The host re-reads its own diff against the task criteria and every reopen reason before recording the result.

Repairs are re-reviewed over the repair delta and affected behavior, reusing conclusions only while inputs remain valid. At integration, new interactions and invalidated conclusions are reviewed rather than automatically re-running every lens.

Loop bounds stop repetitive strategy, not authorized work — never use them to ship defects. Serious disputed defects require evidence inspection or main-model diagnosis: no confidence score can dismiss a failing test, and refuted findings get a corrected independent report rather than a silently edited review.

Resume after interruption

references/recovery.md

Run revisions and immutable artifacts live under the target workspace’s neutral .amaleh directory and must be retained when clearing chat or moving work. This release uses explicit material-transition checkpoints and skill-driven restoration: no native hooks are installed and no native summary replacement is claimed, so run state stays durable even if a final response or compaction hook never occurs.

Invoking the skill again lists the workspace’s runs and resumes the matching one from its last checkpoint, without the old chat. status is the small restoration index: intent, constraints, accepted decisions, current task, unresolved failures, escalation history and next actions.

resume checks worker process liveness on the recorded host and compares accepted workspace fingerprints. Interrupted work becomes blocked for reconciliation, not automatically ready: inspect session logs and actual changes, determine whether pending side effects occurred, then requeue with evidence. A live worker is not stale merely because the chat was cleared, and unlock only removes a demonstrably dead local lock owner.

Credit exhaustion and persistent service failures preserve state; transient API retries are bounded, and once credentials or services are restored the same run resumes. An ambiguous operation result requires checking external reality before replay — this is not a guarantee of exactly-once remote execution.

When a review workspace is missing or unreadable, next and status return reconcile-workspace while independent ready candidates keep progressing. resume preserves original artifacts and revision history, invalidates affected nonrunning verification without changing repair ancestry, and leaves live descendants owned by their workers; their results cannot be accepted until prerequisite evidence is reconciled. The runtime does not restore or relocate a checkout automatically.

Closing the host does not keep a daemon running. A checkpointed run continues on the next invocation, in another session or on another day.

Diagnose, host actions and export

references/runtime.mdreferences/recovery.md

Use diagnose before retrying unexplained failures. It returns current state, pending work, operation timelines, failures, unfinished operations and usage totals, and works even when initialization failed before a valid snapshot existed.

Its health section measures delegation quality — coordinator-authored decisions per task, worker-side Jev calls, delegate versus manual dispatches, host takeovers per task, host-action ceremony and model-family distribution — and emits warnings naming anti-patterns such as coordinator micro-decisions, manual stepping through the worker→review→repair loop, or one task carrying the whole feature. finish reads the same warnings and refuses on them.

host-action records an immutable, local host observation even before run initialization or after an API failure. Kinds are decision, worker, review, edit, check, integration, permission and other; phases are planned, permission-granted, permission-denied, started, completed, failed and skipped. A retry uses a fresh action id, and summaries must never contain credentials. diagnose includes this ledger — the records are host attestations, not proof of tool execution.

host-action input
{ "actionId": "review-attempt-1", "sessionId": "host-session-1",
  "kind": "review", "phase": "planned",
  "summary": "Request fresh review of the current task",
  "taskId": "layout", "next": "Request host network permission" }

diagnostic-export writes a local JSON file under the run’s exports directory and returns its path and content. Its structural allowlist excludes free text, prompts, code, original action/session/task IDs, model names, paths, raw tool output and exception messages; it retains pseudonymous action/session relationships, phase and kind, task status counts and usage summaries. It works without valid run state, and nothing is uploaded automatically — inspect the result before explicitly sharing it.

Costs are separated into OpenRouter response-reported cost for Jev calls and estimated cost for pi worker and reviewer calls. Each pi call is priced from the OpenRouter catalog card that routing recorded before it started, and keeps pi’s local estimate only when no card exists. Neither is a reconciled billing ledger, and missing usage is unknown, not proof of zero cost.

The exact inputs and outputs of every operation named here are in the command reference; the review gate’s place in the run loop is on the workflow page.