Skip to content

We're open source. If actrone-memory has been useful to you, a star on GitHub means a lot to us.

How actrone-memory’s two-tier memory works

Where recent turns live, what goes into long-term memory, how hybrid recall ranks it and how a token budget is split, with numbers you can reproduce.

Matthew Nyirenda(LinkedIn, opens in a new tab)

Founder and CEO, Apocalypse Technologies

5 min read

Rewritten 1 October 2026. This post replaces “How we designed two-tier agent memory: Redis + Qdrant” from April 2026, which described an earlier design. Several of its details no longer matched the library, so it was rewritten against actrone-memory 0.2.1 for Python and 0.1.2 for TypeScript.

The single biggest failure in agents with memory is context. Either the agent re-sends everything on every turn, which is expensive and runs into context limits, or it sends nothing and forgets what it was doing, or it sends a poorly chosen subset and misses the one thing that mattered.

actrone-memory handles this with two tiers and a retrieval pipeline that fits what it returns into a token budget you choose. This post walks through how that works today, in Python 0.2.1 and TypeScript 0.1.2. The two libraries share one model, so everything here applies to both unless it says otherwise.

Two tiers with two jobs

The first tier, L1, holds each session’s recent turns: what the user and the agent said most recently. It keeps the last max_session_turns, 50 by default, and is read on every call whatever the query, so the agent always knows where the conversation stands.

The second tier, L2, holds long-term memories: facts you write with inject_memory or injectMemory, facts the opt-in extractor distils from conversations and, in Python, session summaries. Each one is embedded and indexed for search. Recent turns are not copied into L2; they stay in L1.

Each call to retrieve_context or retrieveContext fetches from both tiers at the same time, so you wait for the slower fetch rather than the sum of the two. The result says what it used:

typescript
import { MemoryManager } from 'actrone-memory'

// In-memory + local embedder by default: no Redis/Qdrant required to start
const memory = await MemoryManager.create()

await memory.storeTurn(
  'research-agent',
  'session-42',
  'Summarise Q4 earnings for AAPL',
  'Apple reported revenue of $119.6B in Q4 2024, up 6% YoY...',
)

const context = await memory.retrieveContext(
  'research-agent',
  'session-42',
  'Apple revenue Q4',
  2000,
)

// context.recentTurns          : recent session turns (L1)
// context.episodicMemories     : semantically relevant long-term memories (L2)
// context.totalTokensUsed      : tokens consumed across both tiers
// context.retrievalDurationMs  : retrieval latency in milliseconds

Nothing to run by default

With no arguments, MemoryManager.create() backs both tiers with an in-memory store and a local embedder. There is no database to start and no API key to get. The store lives in your process, so it is gone when the process exits: fine for development and single-process apps, and the wrong choice for production.

For production, use the stores you already run. Redis, a Redis-protocol server such as Valkey, or Postgres can hold recent turns. Qdrant, or Postgres with pgvector, can hold long-term memories.

Your application code does not change when you switch. The stores sit behind two small interfaces, L1Store and L2Store, and both libraries ship a conformance suite, so a store you write yourself can prove it behaves like the built-in ones.

How recall ranks long-term memories

By default, L2 ranks candidates with Reciprocal Rank Fusion over three separate rankings: vector similarity, keyword match (BM25) and recency. Each ranking adds 1 / (k + rank) to a memory’s score, so a memory does not need to win on embeddings alone. An exact keyword match, such as an error code or a product name that an embedding blurs, can still bring it to the top.

With hybrid retrieval turned off, ranking falls back to a weighted blend: 0.7 times the cosine similarity plus 0.3 times recency, where recency falls in a straight line from 1 for a memory written now to 0 for one 30 days old. For extra precision you can turn on a cross-encoder reranker, which rescores the top candidates by reading the query and each memory together. It is off by default because it adds latency.

Thresholds belong to the embedder

A memory reaches the context only if its similarity to the query clears a relevance threshold. Similarity scores are not comparable across embedding models: the keyword-only hashing embedder scores relevant text around 0.24, while bge-small scores unrelated text around 0.48. Before our first public release, one fixed threshold of 0.72 applied to every embedder. On a labelled set of 48 relevant and 528 unrelated query and memory pairs, it recalled 4% of the relevant memories with the hashing embedder, 23% with MiniLM and 65% with bge-small.

Each built-in embedder now carries the threshold it was calibrated for on that set:

EmbedderThresholdRecallPrecision
bge-small-en-v1.5 (the Python [onnx] extra, or fastembed in TypeScript)0.6388%84%
all-MiniLM-L6-v2 (the Python [local] extra)0.4088%91%
Hashing, keyword only (no extra installed)0.3042%36%

An embedder with no calibrated value, such as OpenAI’s text-embedding-3-small or your own, starts at 0.72 in Python and 0.7 in TypeScript. Measure it on your own data before you rely on that number.

Splitting the token budget

You pass a token budget with each call. The pipeline gives 35% of it to recent turns and 25% to long-term memories, and leaves the remaining 40% for your system prompt and the current turn. Python names that reserve explicitly, 30% for the system prompt and 10% for the current turn, and checks that the four shares add up to 1.

Each share is filled in priority order and pruned to fit. An unused share is not handed to the other tier, and retrieval never touches the part reserved for your prompt.

Summaries and extracted facts

In Python, after 20 turns by default, a background task summarises the session and writes the summary to L2. With the OpenAI provider configured, gpt-4o-mini writes it; otherwise the summary is extractive, taken from the conversation itself, so no data leaves your machine by default. The raw turns stay where they are, and a per-session cooldown of 300 seconds stops two summaries from racing. The TypeScript library does not summarise.

Fact extraction is opt-in in both languages, because every extraction is a model call. It works with any OpenAI-compatible server, including a model running locally on Ollama, and tags each fact with a sensitivity: none, low, pii or sensitive.

What it costs, measured

We have no production fleet to quote numbers from yet, so here is what you can reproduce. The Python library ships a quality benchmark that runs offline:

bash
pip install actrone-memory
python -m actrone_memory.benchmark

It stores memories for an agent, then asks 14 questions whose right answers are known. With the default in-memory store and the keyword-only hashing embedder, it reports:

MetricResult
recall@51.000
precision@50.200
Mean reciprocal rank (MRR)0.943
Median retrieval time0.27 ms, in memory, on one development machine

Each question has one right answer, so precision@5 cannot rise above 0.2. Fourteen questions is a small set, so treat these as a regression gate: CI fails if recall@5 drops below 0.85. On the same questions, a baseline that returns the newest memories scores recall@5 of 0.571. We have not run a head-to-head against other memory libraries yet, so we make no comparative claims.