AI that remembers what you saidand knows who said it.
Memory for agents and AI tools that keeps your own words apart from web pages, documents and the AI's own conclusions, records when each thing was said, and stays on your machine.
Recall is the Aura 1.60 median on 5,000 records and distinct questions; conditions and limitations are listed below.
Memory you canactually work with.
You do not need to write code to benefit from persistent AI memory. Remy packages it into a desktop agent for real work — with every important action visible and under your control.
Memory that compounds
Remembers decisions, preferences, evidence, and past work across conversations.
Work you can inspect
Trajectory shows messages, model calls, tools, timing, and where a run succeeded or failed.
Workflows, not only chat
Run pipelines, automations, research, and repeatable tasks from one workspace.
A closed Agent Lab
Give Remy a goal and let it plan, create, verify, and return the finished artifacts.
Built to remember.Designed to be governed.
The model stays frozen. Aura changes instead — memory accumulates, beliefs form, patterns emerge. Adaptation is bounded, auditable, and operator-controlled at every step.
4-Level Memory Hierarchy
Working (~days) → Decisions (~1 week) → Domain (~2 weeks) → Identity (months). Each level decays at its own rate per maintenance cycle; records that keep being used are promoted.
Cognitive Pipeline
5-layer reasoning stack: Records → Beliefs → Concepts → Causal Patterns → Policy Hints. Each maintenance cycle builds higher-order understanding from raw memories — zero LLM calls.
SDR Indexing
Native Rust indexing combines Sparse Distributed Representations and Tanimoto similarity. In a 5K-record local test, distinct-query top-10 recall measured 11.61 ms median; cache hits and known-ID lookups are separate measurements.
Belief Formation
Records are automatically grouped into beliefs with competing hypotheses, confidence scores, and conflict detection. Epistemic update phase derives support/conflict from the memory graph.
Explainability & Provenance
explain_recall(), explain_record(), and provenance_chain() expose exactly why a memory was surfaced and how it was derived. Every adaptation stays auditable — operators can inspect, restrict, retract or delete.
Immutable Evidence Lineage
For research evidence, SHA-256 lineage binds each finding to an immutable source revision and its exact byte span. Verification status and answer permission remain independent, explicit gates.
Deterministic Context Capsules
Build namespace-isolated, token-bounded hot context with selection reasons, omission counts, and a stable content hash. Blocked and superseded records are never surfaced.
Governed Adaptation
In the Rust API, capture_experience() and ingest_experience_batch() enable bounded self-adaptation without model retraining. Risk scoring and purge/freeze controls keep autonomous plasticity operator-safe.
Encryption at Rest
ChaCha20-Poly1305 with Argon2id key derivation. Append-only binary storage ensures transactional data integrity and power-loss resilience across edge and cloud.
MCP Ready
Native Model Context Protocol server — works with Claude Desktop, Cursor, VS Code and any MCP client. The Python server has 11 tools, the Rust server 23; an HTTP transport serves Make.com and n8n and requires an API key beyond localhost.
From a signal tolasting cognition.
From input to permanent memory — no LLM calls, no embedding API, no cloud. Pure deterministic computation in Rust.
Input Encoding
Text is converted into a Sparse Distributed Representation (SDR) from xxHash3 n-grams. Deterministic, no neural model needed; embeddings are optional.
Source Label
Every record keeps where it came from: recorded (the user), retrieved (documents, web, tools), inferred or generated (a model). Recall keeps first-hand memory apart from outside text, which is quoted as data, not instructions.
Resonance Search
Candidates come from an inverted index; SDR Tanimoto similarity, n-gram and BM25 scores are fused to rank them. Optional embeddings add meaning-level matches.
Store or Reinforce
An exact duplicate reinforces the existing record instead of creating a new one. Near-duplicates are merged later, during maintenance — never at write time, so negations, versions and dates are not lost.
Consolidation
Maintenance cycles promote records that keep being used (Working → Decisions → Domain → Identity) and build beliefs, concepts, causal patterns and policy hints — without LLM calls.
Decay
Each level decays at its own rate, slowed by use; weak records are archived, identity records never are. Storage is append-only with checksums, resilient to power loss.
No cloud tax.No invented speedups.
Aura runs memory operations locally, without a required LLM or embedding API. Here are measured results for distinct questions, not cache hits or estimates for other products.
| 5K-record local test | Aura 1.60 | SQLite FTS5 lexical index |
|---|---|---|
| New-query top-10 recall, median | 11.61 ms | 1.13 ms |
| New-query top-10 recall, p95 | 18.95 ms | 2.33 ms |
| Store one record, median | 5.66 ms | 14.89 ms |
| Get by known ID, median | 0.0033 ms | 0.0044 ms |
| All annotated evidence found | 36 / 100 questions | 39 / 100 questions |
One local run on Windows 10 / Ryzen 5 5600X: 5,000 LoCoMo dialogue turns, 100 distinct natural-language questions, top-10 results, no repeated-query cache hits. Median and p95 are not service guarantees.
The older Aura 1.58 test at 1,000 different records measured 2.48 ms median for uncached recall and 8.2 µs for a formatted cache hit. Different data and questions mean this is not a version-to-version comparison.
FTS5 is a lexical index, not an equivalent cognitive-memory system. The storage paths do different work, so store times are not a like-for-like durability comparison. The evidence result is a narrow diagnostic slice, not an overall quality ranking. See the benchmark conditions.
Memory at thespeed of thought.
Python SDK works with any LLM framework. Store, recall, done.
4-Level Memory Hierarchy
Memories decay naturally and promote automatically. The cognitive pipeline runs in the background, forming beliefs, concepts, and causal patterns — no LLM calls required.
Current session context. Recent messages, active tasks. Decays quickly unless accessed.
Choices and reasoning. Why you picked X over Y. Promoted from Working on repeated access.
Learned knowledge and code. Project context, technical facts, domain expertise.
Lasting preferences and traits. Decays slowest and is never removed by decay.
Your agent can remember today.
Python 3.9+. Pre-built wheels for Linux, macOS, and Windows. Install locally in one line — no separate memory service required.
MIT License · Rust core
Built in Ukraine
Need commercial licensing or priority support?
View Pricing