How a verdict is issued, and what could flip it.
Everything here is designed so you do not have to trust us: criteria registered before the run, data sealed and hash-anchored, results you can recompute yourself.
A verdict is not binary: it has five layers
The binary — did it win or not — is given away free by the literature. What is charged for is layers 2 to 5.
Do the claim’s artifacts run? runnable / partially / not runnable. Many claims die here.
Under a criterion pre-registered and sealed BEFORE anything runs. Never adjusted by looking at the result.
Percentage gap to the champion baseline, time-to-solution ratio, number of seeds and interval — or their declared absence. If variance was not measured, we say it was not measured.
Sensitivity to size, seeds and hyperparameters, and the strength of the baseline: does it only win against weak ones?
Two or three observable events that would flip the verdict. It is what makes the report age well.
The result in the client’s units. The client’s numbers come from the client or are marked as assumptions; we supply the arithmetic and the sources.
The evidence scale (EL0–EL5)
Adapted from GRADE, the evidence-review standard in medicine: a base level set by the design of the evidence, plus declared factors that raise or lower confidence — never hidden arithmetic, always the reason in writing.
| level | name | what it requires |
|---|---|---|
| EL0 | Claim only | A dated, sourced public claim. No raw data or code downloadable by hash. |
| EL1 | First-party artifacts | Raw data and code published and hash-verifiable: anyone can recompute what the author did. |
| EL2 | Pre-registered & sealed | Criterion registered BEFORE the run, run sealed, the result meets its own criterion, and a classical baseline declared on the same field. Still first-party. |
| EL3 | Independently reproduced | A third party with no involvement or payment from the author reproduced the main result from the published artifacts. Today: empty. Nobody in the registry has it. |
| EL4 | Adversarially tested | EL3 plus surviving a serious published refutation attempt, or being reproduced against the strongest known classical baseline, at parity. |
| EL5 | Verified in production | Independent verification in a real deployment with economic consequence and a named external referee. Today, one: certified randomness. |
The scale measures the strength of the evidence that the advantage is real. It does NOT measure survival over time: that is the status (surviving / contested / eroded / open / negative-selfpublished), which is orthogonal and shown alongside. A claim can be EL2 and eroded; another EL0 and surviving.
The v1.0 rubric is ready to seal and is NOT sealed yet. Until it is, we publish no claim’s level: scoring with a yardstick that can still be changed is exactly what this firm exists not to do.
juez-v1: budget parity
The protocol that makes a comparison comparable. The classical and the quantum side get the same instance, the same compute budget and the same wall-clock, with fixed seeds and declared versions. The baseline is the champion of the class, tuned — not the first thing that compiles. Without declared parity a result is not evidence: it is a demo.
The Two-Layer Rule
Layer A — the sealing infrastructure: mechanical, opinion-free. Anyone may pay for it, vendors included. It notarizes integrity and date, not truth. Layer B — judgment: verdicts, evidence levels, reports, Monitor, Registry, Library. Paid for only by the demand side; never by the party being evaluated, directly or indirectly.
Verify a seal yourself → · The full registry →
Version 1.0 · 1 Sep 2026. Methodology changes are published with dates; verdicts cite the version they were issued under.