Cacheon

The finalized chain loop

The production loop is a non-emitting, restart-safe intake and qualification controller. It does not accept a shell evaluator, keep a JSON scoring ledger, or submit weights after each pass.

One pass

run_pass(...) performs these operations in order:

  1. Bind chain scope. Open the SQLite store against the chain genesis hash and netuid. A database created for another scope is not reused.
  2. Read finalized history. Continue after the durable cursor and reconstruct exact reveal priority from finalized storage and canonical event positions.
  3. Reserve before transport. Persist every arrival in chain order before any fetch. Slow hosting therefore cannot rewrite priority.
  4. Fetch privately. Accept HTTPS only, validate DNS and every redirect, enforce archive limits, extract regular files safely, and rederive the committed hash.
  5. Classify and fingerprint. Parse the target-scoped proposal; a bundle the component parser rejects is refused. Copy identity covers the submitted delta, never the validator's incumbent stack.
  6. Publish immutably. Copy the validated private tree to a content-addressed worker publication and reopen it before use.
  7. Reconcile copies. Compare durable fingerprints in finalized order. This step is separate and idempotent so a crash between publication and copy disposition cannot bypass priority.
  8. Screen and qualify. If a registered arena service was injected, run its staged screens, use the routing-only resident lane where applicable, form a capacity-bounded cohort, execute authoritative resident qualification, and persist outcomes.
  9. Settle retained pairs. Lease economically unblocked, independently reproduced candidates and apply the resulting settlement plan transactionally.

The pass returns counts and dispositions. It never opens a wallet or calls set_weights.

Reservation state machine

The row status is an operational control signal, not merely a progress label:

qualified means two matching PASS qualifications have been retained. The associated settlement candidate then has its own transactional state (pending, leased, and a terminal economic disposition such as crowned, held, neutralized, or discovery_bounty). A reservation can remain qualified while settlement decides its economic outcome; do not infer a crown from the reservation status alone.

Several terminal paths are omitted from the diagram for readability: an operator may explicitly expire sufficiently old inactive work, copy reconciliation can turn a later submission into failed, and bounded retry exhaustion leads to held.

Public CLI: intake only

The supported standalone command is:

cacheon chain-validate \
  --netuid 307 \
  --network "wss://test.chain.opentensor.ai:443" \
  --intake-db chain_intake/intake.sqlite3 \
  --private-root chain_intake/private \
  --publication-root chain_intake/worker \
  --audit-log chain_intake/chain-audit.jsonl \
  --intake-only \
  --once

Remove --once to run continuously; --interval controls the delay between passes.

chain-validate accepts only its declared intake and arena schema. Chain-signing credentials, external evaluator commands, scoring policy, and weight publication belong to separate authorities and must not be added to the validator-loop service.

Without --intake-only, the CLI rejects startup unless its Python caller injects an exact ArenaServiceRegistry and selects a registered --arena-id:

from cacheon.chain.validator_loop import run_validator

run_validator(
    subtensor,
    netuid,
    intake_db="chain_intake/intake.sqlite3",
    private_root="chain_intake/private",
    publication_root="chain_intake/worker",
    audit_log="chain_intake/chain-audit.jsonl",
    arena_registry=registry,   # constructed by reviewed deployment code
    arena_id="production-arena-id",
    intake_only=False,
)

This is an integration boundary, not a copy-paste complete deployment: the repository does not provide the production provider represented by registry.

In daemon mode, run_validator contains pass-level validator faults. It logs the full exception, increases the sleep multiplier up to six times the configured interval, and resets the failure count after a successful pass. The Python API default stops after ten consecutive failures so a supervisor can intervene. Candidate dispositions already committed before the process-level exception remain in SQLite; the loop does not roll the entire pass back as one transaction.

HTTPS intake boundary

The on-chain payload is a canonical JSON object containing schema version, lowercase SHA-256 content hash, and an HTTPS URL. It is limited to 1024 UTF-8 bytes.

Production fetch enforces:

  • HTTPS and TLS 1.2 or newer;
  • globally routable resolved addresses;
  • connection to a reviewed address while retaining TLS SNI and hostname checks;
  • validation of every redirect, with at most five redirects;
  • 64 MiB downloaded archive, 256 MiB extracted content, and 4096 members;
  • 16 MiB per regular file, 8 MiB per inspectable source file, and 32 MiB aggregate inspectable content;
  • raw gzip/tar preflight before tarfile materializes metadata, including 64 KiB per PAX/GNU extension header and 1 MiB aggregate extension payload;
  • no symlinks, hardlinks, special files, duplicate/path-conflicting members, or path traversal; and
  • a bounded transfer deadline.

file:// exists only behind explicit test helpers. It is not accepted by chain-submit, payload decoding, or the production fetch function. Plain HTTP is not accepted at all.

Private and worker storage

The fetch root is validator-owned private storage. The code requires owner-private directories and files and never mounts this mutable intake tree directly into a worker.

Publication creates a separate carrier:

  • all bytes are proven to participate in the committed identity;
  • files are sealed read-only and directories are non-writable;
  • the destination address is content-derived; and
  • reopening independently rederives the publication and content hash.

Treat both roots as operational data, not as interchangeable caches.

Redacted journal and private recovery snapshots

When --audit-log is configured (the CLI default), each completed pass appends and fsyncs one canonical JSONL record. The record contains finalized position, content-derived reservation and receipt digests, counts, and bounded disposition classes. It deliberately omits URLs, hotkeys, candidate bytes, exception messages, wallets, credentials, and ambient environment. A fault record retains only the exception type and consecutive-failure count.

SQLite remains the state machine. The journal is supplementary operational chronology: an audit append failure is reported but cannot cause the loop to replay a pass whose SQLite transitions already committed.

Operator reservation diagnostics

The CPU validator can answer a miner's status question without stopping intake or opening a second writable controller:

cacheon chain-reservation-status \
  --intake-db /srv/cacheon/state/intake.sqlite3 \
  --audit-log /srv/cacheon/state/chain-audit.jsonl \
  --reservation-id <64-HEX-RESERVATION-ID>

Use --content-hash or --miner-hotkey only when it identifies exactly one retained row. Ambiguous selectors are refused. --json emits the same privacy-safe record for a support tool. Both output forms omit proposal URLs, private filesystem roots, and raw exception messages.

Read the result in this order:

  1. arrival_authority is the finalized-chain ordering fact. The order key is block, event index, event subindex, hotkey, and content hash.
  2. queue is present only for queued or active work. A numeric position ranks actual selectable work: reproduction first, then primary, with finalized order inside each class. An active screen or qualification has no queue position. A durable remote lease is leased, has no queue position, and is excluded from the remaining waiting depth; evaluation_lease supplies its stage, generation, cohort position, and expiry block without exposing the private worker owner. Promoted qualification work can be an indivisible retry group or bounded cohort, so it is not assigned a misleading single-row rank.
  3. A screen reject is explained by the terminal typed stage, grade, and evidence digest in screens. A qualification outcome is explained by its persisted PASS/FAIL/NO_DECISION, reason code, report or failure digest, and typed attempt reference when one was retained.
  4. Attribution is derived only from a persisted typed decision. A row whose status is failed but which has no persisted FAIL remains unattributed; status text alone is never used to blame candidate code.
  5. evidence_limitations is mandatory support context. In particular, qualification_failure_retained_by_digest_only means the current schema retained an infrastructure failure product's digest but not a reopenable artifact reference. Report it as validator NO_DECISION, quote the digest, and do not guess at a more specific cause.

The last case is a real current limitation: this command cannot reconstruct bytes that were never durably referenced. Private evaluator logs must therefore remain under the deployment's log-retention policy until typed failure-artifact retention and archive indexing cover every infrastructure path. The command makes that gap visible; it does not claim the gap is closed.

For support across every retained submission by one miner, use chain-miner-report --miner-hotkey <HOTKEY>. It composes the same privacy-safe reservation facts with duplicate-replay history and prints the stated cause and next action for each row. It does not infer candidate blame from status text and cannot recover failure bytes that were retained only by digest. See the CLI reference.

The private recovery mirror is also outside the live transaction path:

cacheon chain-snapshot \
  --intake-db chain_intake/intake.sqlite3 \
  --audit-log chain_intake/chain-audit.jsonl \
  --object-store-bucket <PRIVATE_BUCKET> \
  --object-store-endpoint <S3_COMPATIBLE_ENDPOINT> \
  --sealed-input qualification-inputs=/srv/cacheon/sealed-inputs

cacheon chain-snapshot-verify \
  --manifest-key <PRINTED_MANIFEST_KEY> \
  --object-store-bucket <PRIVATE_BUCKET> \
  --object-store-endpoint <S3_COMPATIBLE_ENDPOINT>

chain-snapshot takes a consistent online SQLite backup, checks integrity and foreign keys, discovers database-referenced immutable worker publications and retained settlement qualification artifacts, and adds only explicitly named sealed inputs. Every object is digest-addressed, bounded during download, and reopened before the manifest is accepted. Models, OCI images, wallets, credentials, caches, unredacted logs, and unrelated evidence roots are not auto-discovered.

chain-snapshot-verify restores into a fresh private staging root (temporary by default), rechecks SQLite, publication receipts and hashes, evidence references, and the redacted journal, and emits a restore map when a retained --restore-root is requested. It never overlays live state. Schedule snapshots separately from validator passes so remote object-store failure cannot enter the SQLite controller's commit path. Protect the archive with a private bucket or policy-isolated private prefix and exercise restore regularly.

Evaluation-lease ownership

FinalizedIntakeStore owns durable evaluation-lease state. FIFO selection order, reproduction priority, qualification cohort formation, lease generations, heartbeat compare-and-swap, expiry, and infrastructure release are store policy. No other module implements a second lease state machine, and deployment tooling must not mutate lease rows with raw SQL.

cacheon chain-evaluation-lease is the tracked one-shot operator adapter over that API: preview, claim, heartbeat, and infrastructure release, with all authority coming from one sealed owner-controlled config file. It is not an evaluation worker, daemon, or scheduler; see the CLI reference.

Site orchestration — launch wrappers, tmux composition, endpoints, wallets, exact filesystem paths, and sealed production configs — stays in the private deployment tree outside this repository. Tracked code owns lease semantics; private operations own only identities and launch composition.

Remote worker transport

Remote execution of leased work is a durable-spool transport, not a second evaluation authority. chain/remote_worker_registration.py binds one worker epoch — endpoint, pinned host keys, commissioned READY receipt, worker readiness, physical lane, interpreter, and shared credential — under one semantic digest. The immutable worker carrier places the validator's .cacheon-native-artifact.json receipt beside the miner's committed source; bundle_hash.committed_content_hash is the one canonical rehash of a receipt-bearing carrier back to the chain-committed identity, excluding exactly that top-level receipt and nothing else. chain/remote_worker_spool.py owns the sealed request/result carriers and their verification; chain/ssh_worker_transport.py shuttles them over host-key-pinned SSH and implements the authenticated transport the remote evaluation dispatcher uses for both screen and qualification; chain/remote_worker_pod_service.py supervises one persistent pod adapter per epoch and parks the epoch on its first command-level adapter failure rather than restarting into an unproven resident model. Transport, pod, and adapter failures surface as infrastructure no_decision records that release the durable lease without consuming an evaluation attempt.

The tracked B300 adapter has two closed construction modes. Screen-only mode executes through the commissioned screen deployment and refuses qualification before resident work. Persistent --serve mode may additionally load one digest-exact qualification-capabilities factory; it then constructs the shared commissioned B300 service, resolves each authenticated promoted cohort, and executes remote qualification through the same READY-bound worker and durable continuation store. One-shot mode cannot commission qualification. Deployment wrappers supply every installed path as an explicit argument; active endpoints, credentials, sealed capability bytes, and process composition stay in the private operations tree.

Standing CPU supervisor

python -m cacheon.chain.standing_cpu_supervisor --config <path> is the standing CPU daemon over those pieces. Its sealed, closed, owner-controlled config names the screen-dispatcher config (chain/mainnet_screen_dispatcher.py supplies the config schema and the dispatcher builder) and the recoverable-qualification authorities to compose, and optionally the settlement network. The screen stage reopens the intake-only validator's durable finalized cursor read-only — rejecting scope drift, regression, and hash changes — claims exactly one durable screen lease at a time, and hands the typed request to the authenticated spool transport. The required ArenaService provider slot is filled by a digest-exact remote-only proxy whose execution methods always fail closed. The qualification stage resumes the same durable request across restarts rather than restarting the experiment. Stage faults tear down the constructed authority and rebuild it under bounded exponential backoff; every status change is one canonical-JSON line on stdout.

enable_settlement installs the transactional settlement stage. When enabled, settlement_network must name the finalized-head endpoint used to clock lease and commit; the stage refreshes that clock immediately before each stateful boundary and never opens a wallet. A commission may stage a nonempty settlement_network while the flag remains false so arming is a reviewed one-field change; enabling settlement with an empty endpoint is refused.

enable_weights installs the eval-side weight-offer push stage (chain/standing_weights_stage.py) and requires weights_stage_config, an absolute path to a second sealed, closed, owner-controlled file with schema cacheon-standing-weights-config-v1 and exactly these fields: network (explicit wss:// finalized-head reader), fallback_endpoint (empty or wss://), push_url (http(s) serve-weights offer endpoint), push_credentials (owner-only path to the push credential set), attribution_hotkey, half_life_blocks, discovery_lifetime_blocks, discovery_pool_ppm, refresh_blocks, and burn_hotkey. Every refresh_blocks the stage reads the finalized head and metagraph, reopens the intake store, and pushes the current V1 offer: the real projection whenever an active reward claim, a crowned arena, or an activated composition exists; otherwise the full-pool burn offer to burn_hotkey when that field is set, or the builder's crownless refusal as a stage error when it is empty. The stage never signs; the serve-weights lane owns readback and the follow-weights signer decides what reaches the chain. Naming weights_stage_config while enable_weights is false is refused, as is the reverse.

Durable reservation states

The store makes work and failure class explicit:

ClassStates
Activereserved, fetching, transport_retry, published, screening, promoted, qualifying, reproduction_pending
Terminalfailed, expired, qualified
Operator/retry dispositionheld, no_decision

Default intake policy bounds include queue size, per-hotkey and per-target admission, transport and qualification retries, cohort size, epoch cutoff, and expiry. These are code defaults, not a promise that they suit every deployment; an operator should review them alongside arena capacity.

The default IntakePolicy values are:

BoundDefaultEffect
Epoch / cutoff360 / 30 blocksArrivals in the cutoff tail are admitted into the next epoch
Pending queue256New valid arrivals beyond the bound fail admission deterministically
Per hotkey / epoch16Limits one submitter's intake occupancy
Per target / epoch64Applied after the target is resolved from submitted bytes
Transport / qualification attempts3 / 3Exhaustion produces a retained hold rather than infinite work
Controller cohort8Bounds fetch, screening, and qualification selection per pass
Finalized-block expiry SLA500,000 blocksKeeps queued work for roughly 69 days before automatic stale-state expiry and sets the minimum age for explicit expiry

The expiry SLA is only meaningful against the queue's service rate. The former 10,000-block bound covered only about 51 reservations at the measured service time and expired 178 queued rows without verdicts. The 500,000-block default keeps automatic expiry as a last-resort stale-state bound rather than a normal capacity disposition. Treat any cohort expiring without being reached as a capacity fault, not a miner outcome.

An exact cohort that expired because the validator worker was unavailable can be readmitted with chain-evaluation-lease requeue-expired --authority <SEALED_JSON>. The closed authority binds the reservation IDs, retained-result IDs, and the fixed validator_worker_unavailable reason. The store restores each row to its durable published or promoted lane and grants a fresh finalized-block SLA without deleting history. The bounded refresh budget fails closed; an owner-escalated repeat must be explicit in a newly sealed authority. This is not a generic expiry undo.

Arena capacity is an additional bound. Its queue age/depth, active-screen, active-qualification, cohort, and retry limits are content-bound in the service manifest. Changing either policy changes operational behavior and should be reviewed and recorded; the code defaults are not calibrated economics.

The controller applies the finalized-block SLA on every pass, including retained-only passes, and inside intake and settlement transactions that depend on unresolved priority. Eligible reserved, transport_retry, published, promoted, reproduction_pending, held, and no_decision rows expire automatically when their arrival or retained-progress block reaches the bound. In-flight fetching, screening, and qualifying rows are not aged out underneath active work. A first retained PASS records a fresh finalized progress block and starts a full bounded reproduction window from that block. Legacy retained evidence with an unknown progress block, including the dedicated schema-3 migration hold, remains fail closed for explicit operator disposition. This prevents slow reproduction from losing its complete SLA while preventing one old PASS from becoming a permanent priority veto.

Eval-cost admission policy

Eval-cost admission is deliberately separate from the shared IntakePolicy used by screen and evaluation-lease services. chain-validate --eval-cost-tao-rao controls the required transfer_keep_alive amount and defaults to 0 (off). Quote TTL defaults to 300 blocks and the payment-to-reveal window defaults to 7,200 blocks.

A v2 reveal may attach a payment pointer. When the gate is enabled, intake rebuilds the remark from that reveal's hotkey, content hash, and netuid; only that triple can spend the pointer. The paying coldkey is not the claimant. When the gate is disabled, the pointer is ignored for payment accounting: unverified coordinates are neither consumed nor allowed to pre-claim a future payment.

Byte-identical resubmissions replay their prior verdict before any lease is claimed: a bundle whose exact content hash already reached a terminal FAIL under the exact current arena service digest inherits that FAIL (reason duplicate_of:<reservation>:<original reason>) and costs neither a screen nor a qualification. A prior PASS is never replayed — settlement requires an independently bound PASS pair, so a resubmitted winner queues for a real evaluation. Any changed byte, or any change to the arena, produces a fresh evaluation.

An operator can grant one artificial make-good with chain-eval-cost-credit. The oldest unspent credit for that hotkey admits one otherwise-unpaid reveal and is consumed inside the same admission transaction. Other admission failures leave it unspent, and a reveal that cites a payment pointer must pass ordinary payment verification instead of falling back to a credit. Credit grants are private validator mutations with an audit note, not a miner-controlled payment token; see the CLI reference.

Verdict and retry semantics

The controller maps failures according to where authority was lost:

Point of failureStored dispositionRetry behavior
Invalid chain payload, unpaid or invalid eval-cost payment, unsafe archive, content-hash mismatch, malformed proposalfailed / FAILNone; attributable intake failure
Eval-cost payment lookup RPC/decode blippass aborted; cursor unchangedRetry the pass; do not fail the miner
Transient HTTPS/DNS or immutable-publication storage faulttransport_retry / NO_DECISIONRetry until the transport budget, then held
Static/build/ABI/graph/serving screen FAILfailed / FAILNone under that screen authority
Screen timeout or inconclusive evidenceRetry in the same primary or reproduction laneArena screen budget decides retry versus hold
Qualification plan/runner/raw-speed failure affecting a registered cohortNO_DECISION for every member plus a persisted bisection planCohort halves are retried to isolate poisoning without assigning losses
Per-candidate post-attempt NO_DECISIONRetained report plus one-candidate requeueRetry in primary or reproduction lane
First complete PASSreproduction_pending; no settlement candidate yetFresh screen and qualification required
Second matching complete PASSqualified; paired candidate becomes settlement-pendingSettlement leases it when earlier economic blockers clear

The qualification retry counter counts retained qualification dispositions. The screen counter counts retained screen attempts. Restarting the service does not reset either. Likewise, changing a reason string or moving files does not create a fresh economic identity.

Restart behavior

On restart, the store does not pretend interrupted work completed:

  • interrupted fetch or qualification becomes held with NO_DECISION;
  • an interrupted screen returns to the appropriate retry lane; and
  • an expired settlement lease returns to pending with a new generation.

The finalized cursor, reservation identities, and immutable publications make repeated passes idempotent. Validator/storage faults should produce retry or hold, not a miner loss. A supervisor can restart the loop, but must not delete or hand-edit the database to “unstick” it.

Recovery is intentionally conservative:

  • fetching and qualifying become held with NO_DECISION, because the controller cannot prove what completed outside the transaction;
  • screening returns to published or reproduction_pending with a retry disposition, preserving which lane was interrupted; and
  • a leased settlement candidate returns to pending, clears its lease, and increments the generation so a stale worker cannot commit it later.

A hold is not self-healing. Diagnose the retained reason, repair the authority, and use the reviewed release/requeue API appropriate to the deployment. The store's release_hold(...) appends an operator reason and chooses the lane from retained publication and reproduction evidence; it does not erase prior attempts. No public CLI wraps reservation-hold release, so deployment tooling must expose it under its own access controls and audit trail.

Archive an exact schema-3 migration hold

One legacy database shape can retain a single-PASS schema-3 candidate that cannot satisfy the current two-PASS parser. It has a dedicated terminal operation:

cacheon chain-archive-schema3-hold \
  --netuid <NETUID> \
  --network <NETWORK_OR_WSS_URL> \
  --intake-db chain_intake/intake.sqlite3 \
  --reservation-id <RESERVATION_ID> \
  --reason "reviewed migration reason"

The command constructs no wallet. It accepts only the exact migration hold, records the current finalized height and bounded operator reason, preserves candidate and qualification bytes, removes the permanent queue veto, and can never release or crown the evidence. Generic expiry and hold release are not substitutes.

Incident playbook

AlertImmediate containmentSafe recovery criterion
Finalized cursor regression or changed hashStop the controller; preserve DB and endpoint logsChain endpoint/finality authority is understood; never overwrite the cursor
“another intake controller owns this database”Find the legitimate owner; do not remove .lockExactly one live controller/signer window owns the DB
Repeated transport retryPreserve URL, DNS, TLS, redirect, and archive evidenceSame committed bytes can be fetched within policy, or work remains held
Publication faultStop worker consumption of the affected addressStorage ownership/modes and independent reopen pass
Growing queue ageStop new operational expansion; inspect screen/qualification capacityRegistered capacity and hardware can drain finalized order without reordering
Cohort-wide NO_DECISIONPreserve failure digest and retry groupsBisection or infrastructure repair completes under the same frozen authority
Evidence root unavailableBlock settlement and weightsExact referenced artifacts reopen; rebuilding “equivalent” JSON is insufficient
Repeated pass exceptionsLet the bounded loop exit and quarantine the host if neededRoot cause fixed; one --once pass succeeds before daemon restart

Operations checklist

  • Put the database, private root, and publication root on durable local storage.
  • Schedule chain-snapshot against a private object-store namespace and run chain-snapshot-verify after every backup; copying only the main file while WAL writes are active is not a valid backup.
  • Test a retained fresh-root restore before relying on the archive, and keep bucket access policy, encryption, versioning/object lock, lifecycle, and capacity alerts under operator review.
  • Alert on growing held, no_decision, transport retry, and queue-age counts.
  • Monitor disk and inode use in both private and immutable publication roots.
  • Run cacheon chain-compat after changing the Bittensor SDK.
  • Keep coldkeys off this host. Intake needs no wallet; the separate weight signer uses only the configured validator hotkey.
  • Coordinate signer access between passes because the SQLite authority is single-owner.
  • Retain service manifest, policy, logs, screen receipts, qualification artifacts, and software/image digests long enough to explain every standing claim.

Continue with Arena service and Settlement and weights.

Source anchors

On this page