Implementation and launch
Section 20 of 22
The handbook defines required behavior. Check the running code against these requirements during implementation and each release audit. Record the code and configuration that enforce each requirement.
Required implementation
| Workstream | Required deliverable | Dependency it satisfies |
|---|---|---|
| Public contribution path | Authoritative merge history, contributor access, source and corpus separation, licensing, and conflict-safe synchronization where needed. | Outside work can enter U01 without bypassing acceptance or creating an implicit payment. |
| Records and controller | Versioned records, authenticated events, rule checks, job deduplication, deadlines, and resumable execution. | U, T, B, R, A, G, H, and payment flows preserve their state across worker failure. |
| Evaluation | Prescribed deterministic and live tests, reproducible inputs, authenticated reports, budgeted adversarial search, and coverage reporting. | Required evidence can control progression without a universal subjective approval. |
| Council qualification | Benchmarks, assignment scopes, agreement rules, injection tests, replacement, and bounded appeals. | Conditional judgment has a qualified authority and a stopping point. |
| Test deployment | Reachable candidate, immutable identity, funded resources, and test-only credentials throughout readers and contracts. | Members can examine U11 without admitting unapproved production evidence. |
| Production authorization | Scoped controllers, transaction enforcement, confirmed activation, valid fallbacks, and coherent client and registry checks. | U14 and U15 refer to the same approved release. |
| Credential lifecycle | Domain-specific outputs, source/build/runtime dependencies, old-format treatment, and effective revocation and distrust. | F3 reports validity consistently before and after an incident. |
| Challenges and incidents | Intake, evidence access, blocking scope, guardian permissions, emergency admission, direct notices, and follow-up. | Serious evidence receives protection without waiting for award adjudication. |
| Funding and settlement | Earned-revenue accounting, reservations, reserve checks, due entitlements, bounded conversion, and confirmed payouts. | Paid intake opens only for promises supported by actual funding and tested settlement. |
| Token activation | Implemented 1% covered buy/sell tax, supported payment and conversion routes, pool/burn routing, and net-payout tests. | Token-dependent capabilities activate after those checks. Release-controller rehearsal uses simulated receipts without requiring a token launch. |
| Participation | Stable identities, public work credit, scoped roles, wallet rotation, lost-key recovery, and voting rules if activated. | Account or key changes preserve history and earned claims. | | Operations and recovery | Gas allowances, key replacement, provider recovery, resource limits, retained builds, and reconciliation. | Common failures reach a valid recovery or explicit terminal outcome. | | Participant terms | Open-source license, report scope, funded award schedule, payment eligibility, privacy handling, and dispute terms. | Participants can understand the work and payment promise before accepting it. |
Do not choose a custom Git service, different chain, or new identity system merely because the automation needs an access credential. Test existing hosting and permission mechanisms first. If they cannot satisfy a documented requirement, compare the specific alternatives within a bounded investigation.
Delivery sequence
- Use the rehearsal profile below to define one safety candidate, its mandatory checks, named role keys, public challenge rule, and funded compute allowance. Record the concrete inputs before running it.
- Implement the record controller and promotion controller with the routine executor as the actual transaction caller. Transfer ownership of disposable test registries to that controller. Qualify the initial council assignments against recorded benchmark expectations before granting their test authority. Exercise admission, pause, cancellation, expiry, and replacement before the passing-candidate drill.
- Run the full evaluation and coverage task using the qualified assignments, open a real test challenge route, and complete controller-authorized promotion against test-only records. Measure its cost and retain every failure and retry. Shortened rehearsal clocks prove transition behavior only.
- Grow the corpus, rotate benchmarks, and requalify changed councils. Exercise public and confidential challenges, a coupled weakening, a late attack finding, emergency containment, and a bounded patch. Open public defect intake early when triage capacity exists. Label any contingent unfunded award and budget regression work separately.
- Use measured cost and access results to adopt the founding operating profile. Before outside credentials depend on it, verify the production custody migration, test isolation, full-duration challenge route, incident powers, and the actual founding recovery limit.
- Establish funded quotes, reservations, observation, remedies, attribution disputes, and recipient recovery before paid intake. Funding work can proceed in parallel. Keep fixed obligations covered independently of hypothetical token trading.
- Activate the token tax, token service payments, and token settlement after their integration and liquidity tests pass. Exercise pool/burn routing with actual assets and reserve accounting. Enable further governance and independent recovery after their rules and controllers are adopted and tested.
Token activation can occur earlier if its own prerequisites are ready. This sequence expresses dependencies, not an assertion that a market or sponsor cannot exist yet. Paid bounties, token trading, and weighted governance are unnecessary to prove the ordinary release state machine.
Rehearsal profile
Use this starting configuration for test-only rehearsal. It requires the listed code and authenticated workers. Rehearsal results do not certify a production policy or authorize spending.
| Setting | Proposed rehearsal value | Required binding or completion |
|---|---|---|
| Domain | Safety for typescript.precise.v1 only. | Freeze the selected source revision, policy rules, model configuration, input schema, and corpus in a rehearsal record. A profile name alone is insufficient. |
| Candidate scope | One full candidate against a fixed baseline, using contributor-owned or synthetic inputs. | Include a harmless change and a deliberately defective variant as separate attempts. Never issue accepted customer credentials. |
| Deterministic checks | Run the complete policy-package tests and corpus command, plus the mandatory controller, isolation, identity, and credential drills listed below. | All required outcomes must pass for the passing attempt. A missing test implementation is incomplete work. |
| Live semantic checks | Three prescribed trials per selected case. Select at least one attack and one benign control for every behavior in this candidate's coverage map. | Freeze case identifiers and expectations before results. Require every mandatory trial to meet its expectation for this rehearsal. This small sample is not a production accuracy estimate. |
| Attack search | Twenty prescribed attempts against the scoped changed behavior, within the candidate's cost allowance. | Record all attempts and findings. A failed worker does not count as a completed attempt. |
| Coverage and challenge councils | Three separately assigned evaluator instances, with two agreeing scoped decisions required. | A reproduced mandatory failure blocks regardless of the vote. Invalid outputs and abstentions do not count as agreement. Separate the challenged author from the challenge assignment. |
| Runner admission | One explicitly admitted test runner key for the frozen runner image and configuration. | Exercise valid evidence, wrong target, tampered output, unauthorized key, revoked key, and stale authorization. Live execution evidence must come from the identified run. |
| Challenge eligibility | Any authenticated participant, with C02 qualification for factual and interpretive submissions. | Inject both kinds in the drill and verify their assigned consequences. No contribution score or token minimum. |
| Challenge time | Ten minutes of qualifying access, including any coupled corpus change. | Use only test resources and disposable identifiers. Exercise an outage and a changed-target restart. |
| Execution delay | Two minutes after confirmed U13 registration. | No overlap with the challenge clock. A timer ending cannot override a pause or missing eligibility. |
| Challenge decision and appeal | Ten minutes for each decision, one reassignment, and one ordinary appeal. | An unresolved blocker closes the attempt after those bounds. Public production uses the longer profile below. |
| Transaction recovery | At most three attempts per authorized external effect. | Reconcile unknown outcomes before retrying. An actual failed behavior test is not a transport retry. |
| Gas | On a local development chain, fund the executor with 0.1 test-native units and permit one 0.1-unit top-up from a designated test account. | No real-asset funding is implied. A different chain requires a newly recorded funded gas allowance before the drill. |
| Compute | A maximum of $25 in metered model use per rehearsal attempt, including its coverage and challenge jobs. | Identify an actual approved source before live calls. Reserve existing hosting separately. If required work cannot finish within the cap, record its cost and incomplete state before revising a future attempt. |
| Money flows | Synthetic fee receipts, reservations, payouts, and burn events only. | Test a threshold of 1,000 unreserved test tokens and the 900-plus-250 crossing example. These values do not become launch economics. |
| Governance | Founding admission and recovery, with weighted voting and optional incentives disabled. | Record role assignments and test that a revoked assignment cannot resume a queued job. |
The repository provides policy definitions and policy test commands (packages/agent-review-policy/package.json) to seed the frozen inventory. The package currently defines test and corpus scripts. Include these inputs in the rehearsal and verify the controller and complete live-model evaluation separately.
The mandatory integration list is exact candidate identification, authenticated runner evidence, full-diff coverage, rejected unauthorized promotion, queue-delay enforcement, pause and cancellation, duplicate-event handling, test credential isolation, revocation reaching readers, partial activation recovery, and a bounded emergency patch. Each item needs a concrete test identifier and pass/fail record before the rehearsal can be called complete.
Founding operating profile
Adopt these starting defaults after measured rehearsal. Review the first operating configuration against the results. Use F9 to change a value before use when measured cost or actual access calls for it.
| Setting | Proposed founding value | Reason and activation condition |
|---|---|---|
| Full release window | Forty-eight hours of qualifying access. | Gives participants across time zones two daily opportunities to examine a change. Validate the actual access and response process before adoption. |
| Targeted release window | Twenty-four hours of qualifying access. | Available only under a proven limited-effect classification. |
| Corpus window | Forty-eight hours, including expectation weakenings. | Coupled work uses the longer applicable duration and separate corpus notice. |
| Execution delay | One hour after confirmed U13 queueing. | Gives the cancellation and recovery route a separate opportunity after eligibility. Prequeue overlap remains disabled initially. |
| Challenge decisions | Twenty-four hours per assignment, one replacement assignment, and one ordinary appeal. | Bound a blocking question while preserving a separate reconsideration. Urgent protection does not wait for these clocks. |
| Council agreement | Two of three qualified, separately assigned evaluator instances for scoped judgments. | Use an explicit agreement result. Any independently validated mandatory failure still blocks directly. Shared operation remains disclosed. |
| Eligibility | Public authenticated factual and interpretive submissions under C02. | Does not depend on activating contribution-based voting. |
| Reserve evidence | No more than twenty-four hours old, invalidated sooner by a relevant ledger change. | Unknown or stale coverage disables burning and keeps incoming allocations in the pool. |
| Retry count | Three transport attempts, followed by one authorized replacement attempt where permitted. | Each action also has a recorded time and cost limit. Reconciliation precedes any repeated external effect. |
| Gas allowance | Reserve three times the estimated cost of the candidate's complete authorized transaction set, at its admitted fee ceiling. | The estimate and actual available native assets must be recorded before queueing. Top-ups stay inside the funded operations allocation. |
| Evaluation budget | Fund the measured mandatory evaluation plus the adopted retry, challenge, and recovery allowances before U03. | A numerical amount comes from the rehearsal and accepted operating prices. A spreadsheet ceiling without assets does not authorize work. |
| Conditional compensation | Thirty days from the accepted observable-use milestone for the substantial-work category. | Each funded quote also specifies evidence requirements and a deploy, assess, or cancel deadline. Smaller-work terms remain separate. |
Production acceptance thresholds and the live evaluation corpus must follow coverage evidence and measured behavior. The rehearsal's three trials and twenty attacks do not establish an adequate production policy by themselves. Bind the adopted test list, model identifiers, qualified keys, and funded budgets before enabling the affected operation. Test exact equality, expiry, missing responses, and all deadline branches in the controller.
Keep the following off until their own prerequisites are met: weighted member voting, token-based eligibility, person-uniqueness requirements, contributor stakes, challenge bonds, unfunded fixed payout promises, private-case evaluation without its examiner, token routes without integration tests, and additional credential domains without a qualified policy. Public code, report intake, and test-only rehearsals can operate with their own bounded resources.
Required drills
| Drill | Result needed before reliance |
|---|---|
| Passing candidate | Sufficient tests and an uninterrupted valid member window produce an authorized promotion without an extra universal council approval. |
| Full-diff coverage | A sensitive untested change requires added checks or remains blocked even when the old corpus passes. A targeted label cannot bypass classification. |
| Founding challenge route | An authenticated participant without voting weight can submit both factual and interpretive claims and receive the specified disposition. |
| Late challenge intake | A complete claim received just before the cutoff reaches qualification before U12 can finish. Duplicate spam cannot restart the window, and an unresolved substantive case cannot pass by backlog. |
| Actual promotion caller | The scoped executor queues and executes through the controller. The founder key performs no routine U14 signature. |
| Confidential case | An assigned examiner can inspect and reproduce the committed case, a public challenge receives a reasoned disposition, and loss of required access blocks completion. |
| Changed candidate | A rebase, model change, or changed expectation invalidates affected evidence and receives the required new challenge opportunity. | | Failed test | The controller blocks progression and cannot clear the failure by dropping or selectively rerunning it. | | Challenge access outage | Lost test access stops the qualifying clock. A nominal deadline does not permit promotion. | | Separate appeal | The assigned authority resolves or closes a blocking case without using the original author as its independent judge. | | Production isolation | A test deployment cannot produce a credential accepted as production evidence by supported readers and contracts. | | Containment and patch | Guardian action stops the affected operation, the bounded patch meets emergency evidence requirements, and failed follow-up returns to containment. | | Revocation | Disposable release and credential records demonstrate the actual effect on checking clients and downstream contract consumers. | | Worker and signer outage | Replacement follows existing authority. Coordinated refusal exercises the independent recovery route or yields an explicit unavailable capability. | | Unknown external outcome | Restarted workers reconcile a pending transaction and produce only one payment or deployment effect. | | Conditional compensation | Accepted work remains recorded through delayed deployment, failed combined release, dispute, and payout outage. | | Funding threshold crossing | Inflows route according to the selected crossing and equality rules, burning never consumes reserved assets, and falling below the threshold restores pool funding. | | Payment conversion failure | Bounded conversion stops and invokes the agreed fallback while preserving any amount owed. | | Tax and net settlement | The actual supported token route respects the 1% covered buy/sell tax, fees, conversion limits, and the recipient's stated net or token-quantity promise. | | Remedy adoption | A new automatic promise cannot activate without funded coverage and an intake exposure limit. Previously accepted remedies survive budget exhaustion. | | Evaluation budget | The complete measured run fits its funded allowance or ends as incomplete with its actual costs recorded. Mandatory checks cannot disappear to fit the budget. |
| Rule migration | Governance authorizes preparation before activation, and failed implementation cannot make the new rule operative. |
Operational review
Set a launch review date and assign a coordinator with authority to propose changes. Review missed defects, false rejections, coverage gaps, challenge outcomes, stale candidates, execution delays, infrastructure recovery, model costs, and payment performance. Compare promised service and reward obligations with actual reserves.
Use observed needs to activate later programs or revise settings. A raw release count or arbitrary case count does not establish safety. Keep the overall launch review separate from individual payment observation periods and repair obligations.