Atlas Librarian · Evidence

Technical leadership should be measurable without requiring Speedway to publish the recipe.

This evidence record exists to separate architectural description from proof. Speedway intends to publish versioned results, methodology, limitations and reproducible result metadata for Atlas Librarian while keeping proprietary ranking logic, private prompts, internal thresholds, security topology and other enabling implementation details confidential.

VersionedEvidence records identify the policy or architecture version and the evaluation suite that produced the result.
ScopedA synthetic contract test is labelled as a synthetic contract test. It is not presented as proof of broad scientific superiority.
ProgressiveComparative and external benchmark results will be added only when the test method and repeated evidence justify the claim.

Current evidence record

AMES starts with behavioural contracts before expanding to comparative benchmarks.

AMES—the Atlas Memory Evaluation Suite—is Speedway's evaluation programme for cognitive memory. Its first public evidence record validates six concrete CMI v2 behaviours using synthetic controlled fixtures. This establishes a repeatable baseline for later ablation and comparative testing.

CMI v2 alpha

Policy version: cmi-v2.0-alpha.1. Evaluation status: additive internal path, with the established fast-recall path retained in parallel.

6 / 6 contract checks

The initial AMES Synthetic Cognitive Recall Contracts suite passed all six defined behaviours on 5 September 2026.

Not an external benchmark

This result verifies implementation contracts. It does not establish that Atlas outperforms another framework, model or research system.

Machine-readable ledger

The sanitized evidence record is published as JSON so technical reviewers and automated systems can distinguish the result, scope and limitations.

Read the evidence ledger →

What the first suite verifies

Evidence is attached to observable cognitive behaviour.

The first AMES contracts exercise properties that matter to a persistent enterprise agent rather than measuring simple text recall alone.

Supersession

When a current decision explicitly replaces an older one, the obsolete memory should not win simply because it remains fresh or frequently used.

Governed rules

Authoritative rules remain available as mandatory context under critical-risk recall even when a tight context budget would otherwise favour cheaper material.

Temporal relevance

When the task state is time-sensitive, scheduler memory can become more useful than a generally similar passive fact.

Outcome evidence

Verified success, failure and rollback evidence can influence otherwise similar memory candidates without becoming the sole authority signal.

Continuation

A project restart can elevate handover and continuity memory even when other project facts remain newer or more frequently used.

Context budget

Optional context is kept within a configured Working Memory budget rather than allowing long-term history to expand without bound.

Benchmark roadmap

Future proof will measure effectiveness, efficiency and failure—not only successful demos.

Speedway's evidence programme is designed to grow from deterministic contracts into repeated comparative tests. Metrics will be published only after the measurement pipeline is stable enough to support them.

Task success

Whether remembered knowledge produces the correct decision, tool choice or governed next action—not merely whether the system can quote an old fact.

Memory hygiene

Rates of stale-memory use, superseded-decision errors, contradiction handling and inappropriate cross-scope recall.

Context efficiency

Working Memory size, context reduction, estimated or measured token usage and the cost of a successful task.

Latency

Median and tail recall latency where results can be collected without exposing sensitive infrastructure topology.

Reliability over repeated runs

Single successful demonstrations are insufficient for claims about agent reliability. Repeated-run success and variance will be recorded where meaningful.

Ablation

Where practical, Speedway will remove or disable memory capabilities in controlled tests to measure what segmentation, lifecycle, authority, adaptive recall and other components actually contribute.

Evidence standard

We will publish failures and limitations as part of the technical record.

A useful evidence programme records where the system fails, not only where it succeeds. Atlas releases can therefore document known failure classes and whether later versions reduce them.

Known limitations

Current CMI v2 evidence is narrow and synthetic. Broader graph traversal, long-horizon consolidation and external comparative benchmarks remain under evaluation.

No silent promotion of claims

An internal contract check does not become a comparative benchmark claim simply because it passed. The evidence classification remains attached to the result.

Version before marketing

Technical claims should identify which Atlas generation they describe so later architectural changes do not rewrite the historical record.

Transparency without IP leakage

Benchmark evidence can be public even when the engine remains proprietary.

Speedway can disclose the test objective, version, number of cases or trials, result, limitations and selected methodology while withholding the exact equations, weights, thresholds, routing rules, prompts and private datasets that implement the system.

Public evidence

Version identifiers, test scope, result counts, measurement definitions, limitations, failure classes and sanitized machine-readable records.

Protected implementation

Exact scoring policy, internal decision thresholds, private prompt design, proprietary consolidation strategy, detailed relationship traversal and security-sensitive internals.

Independent evaluation path

Where commercially and legally appropriate, deeper evaluation can be offered to selected partners or researchers under controlled disclosure rather than publishing trade secrets globally.

Atlas Evidence Ledger

Claims should get stronger only when the evidence gets stronger.