Cacheon

Slot catalog

A slot is a validator-owned semantic boundary inside the pinned engine. A contribution supplies an implementation for that boundary; the validator owns the call site, inputs, output allocation, reference, and verification policy.

The registered API contains 11 slots. The registry in cacheon/slots.py is authoritative; print it with python -m cacheon.cli slots.

Registered slots

SlotKindEntry contractCorrectness
activation.silu_and_mulopentry(x, out)allclose
norm.rmsnormopentry(x, weight, out, eps)allclose
attention.sdpablockentry(q, k, v, out, sm_scale, causal)matched ratio ≥ 0.99
attention.decodeblockentry(q, k_cache, v_cache, req_to_token, seq_lens, req_pool_indices, topk_idx, out, sm_scale, block_size)matched ratio ≥ 0.99
attention.msa_block_scoreblockentry(q, index_k, seq_lens, block_size, out)top-8 overlap ≥ 0.875
attention.msa_prefill_block_scoreblockentry(q, paged_index_k, page/sequence metadata, block policy, out_topk)selected-set overlap ≥ 0.90
moe.fused_expertsblockprepare(w13, w2) + entry(x, topk_ids, topk_weights, prepared, out)matched ratio ≥ 0.97
moe.fused_experts_reducecollectiveprepare(w13, w2) + entry(x, topk_ids, topk_weights, prepared, out, group)matched ratio ≥ 0.97
collective.all_reducecollectiveentry(x, out, group)matched ratio ≥ 0.99
collective.ar_residual_rmsnormcollectiveentry(x, residual, weight, eps, out_norm, out_residual, group)matched ratio ≥ 0.99
collective.moe_finalize_ar_rmsnormcollectiveentry(gemm_out, row_map, scales, residual, weight, eps, out_norm, out_residual, group)matched ratio ≥ 0.99

The callable names in a bundle are selected by its manifest; the signatures above describe their semantic argument order. Entries fill validator-allocated outputs and do not return the tensor consumed by the model.

How to read a signature

Take norm.rmsnorm as the smallest example:

entry(x, weight, out, eps)

The validator creates x, weight, and the scalar eps, allocates out, and invokes the selected entry at the registered SGLang seam. The implementation must preserve inputs and fill the declared output in place. Verification computes a trusted high-precision/reference result for the same case and applies the dtype tolerance. A returned tensor, a different allocation, or a hidden change to x does not replace this contract.

Block slots follow the same ownership rule over a wider semantic region. Collective slots add the validator-owned process group; candidate code may use it but may not create a private group or let only some ranks fall back.

Warning — Unavailable in the current MiniMax-M3 mainnet arena norm.rmsnorm cannot execute because the deployed model uses GemmaRMSNorm rather than the registered RMSNorm.forward_cuda callsite. attention.msa_block_score cannot execute because the pinned runtime has no installing decode-side SGLang adapter. Do not pay for or submit either target to this arena. The entries remain registered ABI and verifier contracts; no other target is withdrawn by this notice. See Current MiniMax-M3 availability.

Kinds

op : A narrow single-device operation. It still executes inside an isolated candidate engine during production qualification.

block : A wider compute region that can express algorithmic fusion while preserving a bounded tensor contract.

collective : A distributed boundary that receives the validator-owned process group. Verification must cover the actual world size and all ranks. Once ranks select a collective candidate, a rank-local failure aborts that candidate engine; falling back on only one rank would diverge the collective.

Correctness policies

ModeMeaning
allcloseEvery element is inside dtype-specific absolute and relative tolerance
matched_ratioA registered fraction of elements must be inside tolerance
cosineCosine similarity and, where configured, relative-norm error are bounded
topk_overlapSelection overlap is measured rather than raw score equality

The standard registered tolerances are 0.02/0.02 for bfloat16, 0.01/0.01 for float16, and 1e-5/1e-5 for float32. Attention targets also carry a 0.03 model-level KL reference in the catalog. These component checks are necessary but not sufficient: production quality authority belongs to the complete qualification profile and pristine reference engine.

Correctness is layered deliberately:

  1. the slot verifier checks the registered tensor contract;
  2. graph replay checks that dynamic inputs are refreshed and outputs are actually rewritten;
  3. the current v7 resident B/C schedule (B′ only when needed) or v8 two-process B/C/B′ schedule measures the exact marginal substitution;
  4. registered eager audit A checks sampled slot behavior outside charged reads;
  5. after candidate teardown, pristine T grades sealed trajectories without candidate code present; and
  6. fidelity/resource evidence checks that the candidate did not obtain speed by bypassing required work.

Passing layer one is necessary to debug a contribution, but only the complete registered profile can issue a qualification PASS.

CUDA graph contract

Production scoring keeps CUDA graphs enabled. Applicable graph-safe entries must use static allocation, avoid host synchronization and data-dependent Python control flow, preserve inputs, and fill only the declared outputs. The validator refreshes registered dynamic inputs and poisons outputs between replays before comparing against fresh trusted references.

An implementation that cannot establish graph safety is not credited for work the baseline performs. attention.decode uses paged caches and validator-selected block IDs at SGLang's captured sparse-attend call; an eager warmup alone cannot produce the qualifying execution receipt.

Graph failure examples

  • A candidate caches the first seq_lens value in Python: refreshed replay inputs expose the stale output.
  • A candidate returns a correct tensor but leaves validator-owned out untouched: poisoned-output replay exposes that the serving call path would consume stale storage.
  • A candidate allocates temporary tensors during capture: graph capture or replay policy fails rather than benchmarking an eager fallback.
  • One collective rank selects the candidate while another selects baseline: the candidate engine aborts because mixed-rank continuation would not define one semantic result.

Slot versus target

Slots define execution ABIs. Targets define reward identities. Most targets map one-to-one to slots, but the target catalog also contains the atomic collective.moe_epilogue.v1 target and a composition rule for the two MoE expert targets. See Target catalog.

For the invariant waist behind every slot, read Target and slot contract.

On this page