Mnemon: New Memory Agent Splits Fast and Slow Reasoning, Achieves SOTA Long-context Results
Mnemon introduces a dual-system memory approach for LLMs, retaining raw conversation records and using a fast decision model for judgments and a slower LLM for reasoning. This results in state-of-the-art long-term memory performance with minimal cost growth and efficient evidence management.
Mnemon is a new memory agent architecture designed to provide long-term memory for LLM-based assistants by storing conversations as raw, dated records rather than rewriting them into facts or structured data at write time. This enables greater flexibility and compatibility with any storage system that returns dated records.

Dual-system architecture: Fast and slow memory processing
Mnemon divides memory processing into two systems: System 1, a fast decision model named Jev that makes many small, independent judgments about whether records are relevant, and System 2, a slower LLM that plans searches and composes answers. This mirrors cognitive theories of fast and slow thinking, allocating most workload to the efficient System 1 and reserving System 2 for complex reasoning.
Efficient evidence consolidation and retrieval
A background process periodically consolidates conversation records into topic timelines, value histories, and standing instructions. This allows questions about a whole conversation to access evidence that individual searches might otherwise miss, improving the quality of answers while minimizing context size.
Performance and cost efficiency
Mnemon achieves state-of-the-art accuracy scores of 91.7% on LoCoMo and 83.8% on LongMemEval-S benchmarks using gpt-4.1-mini, with the lowest effective cost index on LoCoMo. Using a reasoning model, Mnemon reaches 92.2% on LoCoMo and 94.4% on LongMemEval-S—matching or exceeding published results. Notably, the cost per question increases only 1.11x when scaling from 100,000 to 10 million tokens of context.
Implications for developers
Developers can consider adopting Mnemon’s dual-system raw record strategy for scalable, efficient, and accurate long-term memory in LLM assistants, especially when handling large histories. Since Mnemon doesn’t require conversion of records at write time and works with any storage that supplies dated records, integrating this approach can enhance performance and reduce operational cost.
