Build
Read the corpus and write a persistent knowledge base. No questions are visible.
Shattered Needle v3.3 / frozen instance
Shattered Needle compares AI models on a simple job: read many documents, organize what they say, and answer questions later.
The documents contain indirect, ambiguous, and contradictory information. No single document contains the full answer. The model must connect facts across documents and remember which sources support them.
These results come from one fixed investigative test. Higher scores are better. Ties are ordered by wall time.
Higher and farther left is better: stronger results at lower cost. Cost is logarithmic; weighted score is linear from 80% to 100%. Muse appears at both contributor and standard commercial pricing; both are eligible for the dotted frontier. Hover or tap a point for run details.
The performance and cost chart could not be loaded. The complete results remain available in the table.
Runs is the number of completed v3.3 tests. Each run has six reading passes followed by one question session without the source documents. API cost uses the provider bill when available and a token-rate equivalent otherwise. Wall time adds the active test time and excludes gaps between sessions.
Flex estimate reprices the same recorded GPT token usage at OpenAI’s 50%-lower Flex rates. These are cost scenarios, not measured Flex runs; Flex may be slower and cache hit rates may differ. Scores and wall times describe the original runs. A dash means no Flex estimate is published. The chart uses API cost.
Across all displayed rows, API cost and wall time values outside ±1 population standard deviation are highlighted: green for lower and red for higher. Cost is calculated in log space and includes both Muse pricing scenarios; wall time uses ordinary arithmetic values. Flex estimates are excluded from the comparison band.
Shattered Needle tests whether a model can turn difficult documents into useful knowledge. The documents include aliases, indirect references, corrections, and reports that disagree.
The questions are hidden while the model reads. Later, a fresh session must answer them from the model's saved knowledge base. It cannot reopen the source documents.
The model can organize its notes however it wants. A program checks the answers and their source references.
The model knows the task type but not the questions during ingestion.
Read the corpus and write a persistent knowledge base. No questions are visible.
End the session, freeze the artifact, and physically remove the corpus.
A fresh session of the same model answers held-out questions from the artifact alone.
| Model | Culprit | Resolution | Condition | ID provenance | Finding provenance |
|---|
All listed results are completed v3.3 tests on the same frozen investigative corpus. Runs is the number of independent completed draws for each model.
Interrupted or failed tests are excluded from the table and chart until they finish and receive a score.
Actual OpenRouter and xAI values are reported bills. Estimated values apply the public standard token rates recorded by the benchmark to captured usage:
input + cache writes + cache reads + (output + separately reported reasoning)
Unmetered subscription runs have no actual billed amount. Their API cost is a pay-per-token equivalent. These values are lower bounds when a provider applies a per-request long-context surcharge that aggregate token totals cannot reconstruct.
GLM subscription runs use one median-input-price OpenRouter provider’s full rate card (Friendli). GPT estimates use first-party OpenAI prices. Rates checked —; see the pricing audit.
The separate Flex column uses first-party OpenAI Flex rates for input, cache writes, cache reads, output, and separately reported reasoning. It assumes the same recorded token usage and short-context tier, not half of a historical bill. Flex processing trades lower token prices for slower responses and occasional resource unavailability; standard-tier fallback calls incur standard rates.
Sources: OpenAI, Anthropic, Kimi, MiniMax, Z.AI, xAI, and OpenRouter.
A private generator starts from target propositions, lowers them through a function-free Horn derivation graph, and realizes only the leaves as natural-language documents. The forward engine verifies solvability and uniqueness. Distractors are certified not to license an alternative derivation, and all corpus files are frozen by SHA-256 manifest.