Skip to content

We're open source. If actrone-memory has been useful to you, a star on GitHub means a lot to us.

Every integration, tested against a real model

We ran all 27 framework integrations through their real frameworks and a small local model. What broke, what we fixed, and how each check works.

Matthew Nyirenda(LinkedIn, opens in a new tab)

Founder and CEO, Apocalypse Technologies

5 min read

actrone-memory ships adapters for 16 Python frameworks and 11 TypeScript ones. Each adapter is glue between our memory manager and a framework that changes on its own schedule.

Until this release, we checked that glue against each framework’s types and against stub models. That catches a renamed parameter. It cannot tell you whether the memory ever reaches the model.

For Python 0.2.1 and TypeScript 0.1.2 we checked both: the code against the real frameworks, then the result against a real model. Both checks found bugs we had shipped.

Checking the code against the real frameworks

Each library’s public Framework Compatibility workflow installs every framework at the oldest and the newest release its supported range allows, then drives the adapter through the real framework. In Python, contract tests run real agents with stub models. In TypeScript, a canary per framework calls the adapter the way that framework’s own API expects, and the compiler checks it.

That turned up six Python adapters the current frameworks rejected or misread:

  • Microsoft Agent Framework: the adapter targeted the pre-1.0 API, which 1.x removed. It now implements the 1.x ContextProvider.
  • LangGraph: ActroneCheckpointer was a plain class, so graph.compile(checkpointer=...) refused it. It is now a real BaseCheckpointSaver.
  • Haystack: the retriever and writer were not registered components, so Pipeline.add_component refused them.
  • LlamaIndex, AutoGen and DSPy: each returned a shape its framework no longer reads, or ignored an input the framework now sends. The Python 0.2.1 changelog lists every fix.

On the TypeScript side, the code was fine and the package metadata was not. npm treats an installed framework outside an optional peer range as an error, ERESOLVE, rather than a warning. Our ranges stopped below the current major versions, so a project on the Vercel AI SDK 7 could not install 0.1.1 at all. 0.1.2 widens them to ai >=5 <8, @mastra/core >=0.10 <2 and @voltagent/core >=0.1.14 <3.

Pinning the floors down exposed two package names that had changed hands. npm’s genkit 1.0.0 to 1.0.3 and agents 0.0.1 and 0.0.2 were published by other projects before Firebase and Cloudflare took those names, so the floors now start at the first real releases.

Checking the result against a real model

A stub model proves the adapter calls the framework correctly. It cannot prove that the memory lands in the prompt, or that a model can use it once it does. So we ran every integration through its real framework against qwen2.5:3b, a 3-billion-parameter model served locally by Ollama.

Every framework gets the same test:

  1. Store one fact a model cannot guess: the payments team’s escalation codeword is PELICAN-42.
  2. Put a recording proxy between the framework and Ollama, so every request is logged.
  3. Ask for the codeword through the framework.

A run passes only when a request the framework sent carries the fact, the model’s reply contains the codeword, and the adapter stored the real exchange. A control run with no memory shows the same model never produces the codeword on its own.

All 29 runs passed: 16 Python frameworks, 11 TypeScript frameworks, and the library on its own in each language. Each framework ran at the newest release our declared range allows.

Python frameworkVersion
LangChain1.6.6
LangGraph1.2.12
CrewAI1.15.23
AutoGen0.7.5
LlamaIndex0.14.25
Haystack3.2.0
DSPy3.4.0
Agno3.0.11
smolagents1.26.0
AWS Strands1.57.1
OpenAI Agents SDK0.22.3
Pydantic AI2.52.0
Claude Agent SDK0.2.159
Semantic Kernel1.36.0
Google ADK2.10.0
Microsoft Agent Framework1.19.0
TypeScript frameworkVersion
Vercel AI SDK7.0.126
LangChain.js1.2.13
LangGraph.js1.4.18
Mastra1.72.0
LlamaIndex.TS0.12.1
OpenAI Agents JS0.18.0
Firebase Genkit1.42.0
VoltAgent2.11.0
Claude Agent SDK0.3.286
Cloudflare Agents, as a Durable Object in workerd0.24.0
Inngest AgentKit0.13.2

What the real model found

Fact extraction came back empty

Fact extraction reads a conversation and stores the durable facts in it, each tagged with a sensitivity. With qwen2.5:3b it came back empty on 12 of 15 real exchanges. The model read the assistant’s reply as part of what to mine, and returned nothing.

Extraction spec 1.1, in both libraries, frames the conversation, says to record only what the user states, asks for a JSON-schema structured output and gives two worked examples: one with facts and one where the answer is none. On the same model it found 14 of 14 expected facts and stored nothing for a thank-you. If a server rejects the schema, the extractor falls back to JSON mode.

A 3B model still has limits. After a general question it sometimes records that the user asked about the topic, and it tags some preferences and feelings as none. Check the sensitivity tags before you rely on them to redact anything.

AutoGen needs one setting for other models

AutoGen refuses a second system message unless the model client declares multiple_system_messages in its model_info. Memory arrives as a system message, so this affects any memory, including AutoGen’s own ListMemory. For a model AutoGen does not already know, such as one served by Ollama, set it yourself. The adapter’s docs now say so.

Agent Framework moved its OpenAI clients

Microsoft Agent Framework 1.19 moved its OpenAI clients into a separate agent-framework-openai package and renamed model_id to model. OpenAIChatClient now speaks the Responses API, so a chat-completions server needs OpenAIChatCompletionClient instead. If an upgrade breaks your imports, that is why.

Type stripping hid a type error

The TypeScript runs used Node’s built-in type stripping, which removes types without checking them. That hid a real bug: OpenAIFactExtractor typed each message role as a plain string, which the openai SDK’s own types reject, so passing a real OpenAI client failed to compile. TypeScript 0.1.2 fixes it, and a type-checked example in CI now passes the real client.

Try it on your own machine

Fact extraction works with any OpenAI-compatible server, so the model can run next to your code. With Ollama running and qwen2.5:3b pulled, this is the whole Python setup:

python
from actrone_memory import MemoryManager, OpenAIFactExtractor


async def memory_with_local_extraction(
    base_url: str = "http://localhost:11434/v1",
) -> MemoryManager:
    """Extract facts with a model on any OpenAI-compatible server."""
    extractor = OpenAIFactExtractor(
        base_url=base_url,
        api_key="ollama",  # Ollama ignores the key; the client needs one
        model="qwen2.5:3b",
    )
    return await MemoryManager.create(extractor=extractor)


async def learn_from_session(
    memory: MemoryManager, agent_id: str, session_id: str
) -> list[str]:
    """Turn the session's recent turns into stored facts."""
    return await memory.extract_memories(agent_id, session_id)

The fact extraction docs show the TypeScript version. To see the framework checks, open the Framework Compatibility workflow in actrone-memory-py or actrone-memory-ts.

What we have not done yet

The live-model run needs a model server, so it is not part of the libraries’ public CI yet. We ran it ourselves for these two releases. Every result above also comes from one small model, and a different model can fail in a different way. If an adapter misbehaves with yours, open an issue and name the model.