YUPAY the governed multi-model audit

YUPAY (Quechua: to count / to reckon / to audit) is our governed multi-model audit harness. We adopt the audit methodology from Kilo Code / André Lindenberg ("We Audited the Same Codebase…"): give every model the same task, then score issues-found · tokens · cost · latency per model. We run it over our own governed open models and emit one DSSE-signed comparison receipt with a Restraint verdict. 0 CDN.

No M3 weights, no M3 derivative. MiniMax M3 is EXCLUDED-BY-DOCTRINE — its open-weight license restricts military/defense use and MiniMax is PRC-based; SZL demos at Defense Unicorns Warhacker. M3 appears below only as a non-participating reference row, never run.

Participating models
scored on identical task
Value pick · lowest cost/issue
MODELED cost basis
Thorough pick · most issues
single most-thorough pass
Signed comparison
DSSE · the governed difference

Audit comparison (issues / tokens / cost / latency per model)

ModelLabelIssuesRecall TokensCost (USD)Cost/issueLatency
Run the audit to see the governed comparison…

Signed comparison receipt DSSE + Restraint verdict

Run the audit to see the DSSE-signed comparison receipt + Restraint verdict…

locked theorems = 8 {F1,F4,F7,F11,F12,F18,F19,F22} @ kernel c7c0ba17 · Λ = Conjecture 1 · Khipu = Conjecture 2 · SLSA L1 honest / L2·L3 roadmap · receipts: DSSE ECDSA-P256-SHA256 · 0 CDN · trust ceiling < 1.0 (the recommendation is bounded, never absolute).
Honest labels: no key is wired in this Space, so rows are MODELED (issues from the published Kilo benchmark; cost = MODELED tokens × published per-token rates, cited). A row is MEASURED only when a real run happens in-process; M3 is EXCLUDED-BY-DOCTRINE, never run. We never fabricate a benchmark.
Attribution: Kilo Code / André Lindenberg audit methodology (blog.kilo.ai) + MiniMax Sparse Attention paper (huggingface.co/papers/2606.13392) as INSPIRATION. YUPAY is our own implementation over our own open models — no M3 weights, no M3 derivative. See NOTICES.md.

🌐Spaces