Shattered Needle

Shattered Needle v3.3 / frozen instance

Shattered Needle

Shattered Needle compares AI models on a simple job: read many documents, organize what they say, and answer questions later.

The documents contain indirect, ambiguous, and contradictory information. No single document contains the full answer. The model must connect facts across documents and remember which sources support them.

Frozen Mechanical grading Same-model WEAVE and QUERY Results updated
3corpora
1,244documents
~250kcorpus tokens
177held-out questions

Results

These results come from one fixed investigative test. Higher scores are better. Ties are ordered by wall time.

Performance vs. API cost

Higher and farther left is better: stronger results at lower cost. Cost is logarithmic; weighted score is linear from 80% to 100%. Muse appears at both contributor and standard commercial pricing; both are eligible for the dotted frontier. Hover or tap a point for run details.

The performance and cost chart could not be loaded. The complete results remain available in the table.

Runs is the number of completed v3.3 tests. Each run has six reading passes followed by one question session without the source documents. API cost uses the provider bill when available and a token-rate equivalent otherwise. Wall time adds the active test time and excludes gaps between sessions.

Flex estimate reprices the same recorded GPT token usage at OpenAI’s 50%-lower Flex rates. These are cost scenarios, not measured Flex runs; Flex may be slower and cache hit rates may differ. Scores and wall times describe the original runs. A dash means no Flex estimate is published. The chart uses API cost.

Across all displayed rows, API cost and wall time values outside ±1 population standard deviation are highlighted: green for lower and red for higher. Cost is calculated in log space and includes both Muse pricing scenarios; wall time uses ordinary arithmetic values. Flex estimates are excluded from the comparison band.

Benchmark

What it measures

Shattered Needle tests whether a model can turn difficult documents into useful knowledge. The documents include aliases, indirect references, corrections, and reports that disagree.

The questions are hidden while the model reads. Later, a fresh session must answer them from the model's saved knowledge base. It cannot reopen the source documents.

The model can organize its notes however it wants. A program checks the answers and their source references.

  • Release Shattered Needle v3.3
  • Corpora investigative, procedural, narrative
  • Disclosure isolated multi-session batches
  • Grading deterministic, including provenance

Protocol

The model knows the task type but not the questions during ingestion.

01 / WEAVE

Build

Read the corpus and write a persistent knowledge base. No questions are visible.

02 / FREEZE

Remove sources

End the session, freeze the artifact, and physically remove the corpus.

03 / QUERY

Recover

A fresh session of the same model answers held-out questions from the artifact alone.

Technical notes

Full score breakdown
ModelCulpritResolutionConditionID provenanceFinding provenance
Run comparability

All listed results are completed v3.3 tests on the same frozen investigative corpus. Runs is the number of independent completed draws for each model.

Failed and interrupted runs

Interrupted or failed tests are excluded from the table and chart until they finish and receive a score.

Pricing and token accounting

Actual OpenRouter and xAI values are reported bills. Estimated values apply the public standard token rates recorded by the benchmark to captured usage:

input + cache writes + cache reads + (output + separately reported reasoning)

Unmetered subscription runs have no actual billed amount. Their API cost is a pay-per-token equivalent. These values are lower bounds when a provider applies a per-request long-context surcharge that aggregate token totals cannot reconstruct.

GLM subscription runs use one median-input-price OpenRouter provider’s full rate card (Friendli). GPT estimates use first-party OpenAI prices. Rates checked —; see the pricing audit.

The separate Flex column uses first-party OpenAI Flex rates for input, cache writes, cache reads, output, and separately reported reasoning. It assumes the same recorded token usage and short-context tier, not half of a historical bill. Flex processing trades lower token prices for slower responses and occasional resource unavailability; standard-tier fallback calls incur standard rates.

Sources: OpenAI, Anthropic, Kimi, MiniMax, Z.AI, xAI, and OpenRouter.

Generation and certification

A private generator starts from target propositions, lowers them through a function-free Horn derivation graph, and realizes only the leaves as natural-language documents. The forward engine verifies solvability and uniqueness. Distractors are certified not to license an alternative derivation, and all corpus files are frozen by SHA-256 manifest.

InterpretationThis is one fixed investigative test, not an overall model ranking.