Atlas Librarian · Evidence
Technical leadership should be measurable without requiring Speedway to publish the recipe.
This evidence record exists to separate architectural description from proof. Speedway intends to publish versioned results, methodology, limitations and reproducible result metadata for Atlas Librarian while keeping proprietary ranking logic, private prompts, internal thresholds, security topology and other enabling implementation details confidential.
Current evidence record
AMES starts with behavioural contracts before expanding to comparative benchmarks.
AMES—the Atlas Memory Evaluation Suite—is Speedway's evaluation programme for cognitive memory. Its first public evidence record validates six concrete CMI v2 behaviours using synthetic controlled fixtures. This establishes a repeatable baseline for later ablation and comparative testing.
Policy version: cmi-v2.0-alpha.1. Evaluation status: additive internal path, with the established fast-recall path retained in parallel.
The initial AMES Synthetic Cognitive Recall Contracts suite passed all six defined behaviours on 5 September 2026.
This result verifies implementation contracts. It does not establish that Atlas outperforms another framework, model or research system.
The sanitized evidence record is published as JSON so technical reviewers and automated systems can distinguish the result, scope and limitations.
What the first suite verifies
Evidence is attached to observable cognitive behaviour.
The first AMES contracts exercise properties that matter to a persistent enterprise agent rather than measuring simple text recall alone.
When a current decision explicitly replaces an older one, the obsolete memory should not win simply because it remains fresh or frequently used.
Authoritative rules remain available as mandatory context under critical-risk recall even when a tight context budget would otherwise favour cheaper material.
When the task state is time-sensitive, scheduler memory can become more useful than a generally similar passive fact.
Verified success, failure and rollback evidence can influence otherwise similar memory candidates without becoming the sole authority signal.
A project restart can elevate handover and continuity memory even when other project facts remain newer or more frequently used.
Optional context is kept within a configured Working Memory budget rather than allowing long-term history to expand without bound.
Benchmark roadmap
Future proof will measure effectiveness, efficiency and failure—not only successful demos.
Speedway's evidence programme is designed to grow from deterministic contracts into repeated comparative tests. Metrics will be published only after the measurement pipeline is stable enough to support them.
Whether remembered knowledge produces the correct decision, tool choice or governed next action—not merely whether the system can quote an old fact.
Rates of stale-memory use, superseded-decision errors, contradiction handling and inappropriate cross-scope recall.
Working Memory size, context reduction, estimated or measured token usage and the cost of a successful task.
Median and tail recall latency where results can be collected without exposing sensitive infrastructure topology.
Single successful demonstrations are insufficient for claims about agent reliability. Repeated-run success and variance will be recorded where meaningful.
Where practical, Speedway will remove or disable memory capabilities in controlled tests to measure what segmentation, lifecycle, authority, adaptive recall and other components actually contribute.
Evidence standard
We will publish failures and limitations as part of the technical record.
A useful evidence programme records where the system fails, not only where it succeeds. Atlas releases can therefore document known failure classes and whether later versions reduce them.
Current CMI v2 evidence is narrow and synthetic. Broader graph traversal, long-horizon consolidation and external comparative benchmarks remain under evaluation.
An internal contract check does not become a comparative benchmark claim simply because it passed. The evidence classification remains attached to the result.
Technical claims should identify which Atlas generation they describe so later architectural changes do not rewrite the historical record.
Transparency without IP leakage
Benchmark evidence can be public even when the engine remains proprietary.
Speedway can disclose the test objective, version, number of cases or trials, result, limitations and selected methodology while withholding the exact equations, weights, thresholds, routing rules, prompts and private datasets that implement the system.
Version identifiers, test scope, result counts, measurement definitions, limitations, failure classes and sanitized machine-readable records.
Exact scoring policy, internal decision thresholds, private prompt design, proprietary consolidation strategy, detailed relationship traversal and security-sensitive internals.
Where commercially and legally appropriate, deeper evaluation can be offered to selected partners or researchers under controlled disclosure rather than publishing trade secrets globally.
Atlas Evidence Ledger