Hi LangGraph team,
I’d like to propose a BaseStore implementation that adds per-namespace
capacity-bounded warm memory in front of a vector or durable store. Before
opening any PR I’m following the contribution guidance and posting the
proposal here so the abstraction framing can be reviewed by maintainers.
Repo: GitHub - vsingh45/WarmMemory: Capacity-bounded warm memory for LLM agents with LangGraph BaseStore integration, embeddings-based scoring, and comparative benchmarks. · GitHub
Package: warm-memory on PyPI — pip install warm-memory[langgraph] (Python 3.11+, MIT licence)
Full proposal document (with the four required sections — Abstraction /
Usage Pattern / Performance & Scaling / Integration with Existing Tools):
docs/community/proposal.md
TL;DR for triage
- Abstraction:
BaseStoreimplementation. Not aCheckpoint. Not a
Stateextension. Implementsbatch/abatchand conforms to the
standard filter-operator dialect. - Novel choice: per-namespace capacity-bounded eviction. Each top-level
namespace gets its own bounded warm buffer; A’s writes never evict B’s
memory. Bounds total memory atO(namespaces × capacity). - Use case: warm tier in front of a vector store (
InMemoryStore,
PostgresStore, your own). On a synthetic 12-turn workload this
eliminates ~50% of vector-store retrievals and ships the smallest prompt
while maintaining accuracy betweenvector-onlyandfull-history. - Integration: orthogonal to
langgraph-checkpoint*. Mirrors the
langgraph-checkpoint-postgrespackaging conventions if maintainers
prefer a separatelanggraph-store-warmdistribution.
Minimal reproducible example
Runs with no API keys — FakeListChatModel + the default
KeywordImportanceScorer:
examples/minimal_langgraph_warm_memory.py
from langchain_core.language_models.fake_chat_models import FakeListChatModel
from warm_memory.langgraph import WarmStore, build_warm_memory_agent
store = WarmStore(capacity=8)
agent = build_warm_memory_agent(
model=FakeListChatModel(responses=[
"Noted: you prefer concise answers.",
"Yes, you asked for concise answers earlier.",
]),
store=store,
)
agent.invoke({"query": "I prefer concise answers.", "namespace": ("alice",)})
result = agent.invoke({
"query": "What response style did I ask for?",
"namespace": ("alice",),
})
print(result["recalled"])
# [{'key': 'exchange-1', 'score': 0.55..., 'value': {'user': '...', 'assistant': '...'}}]
The package ships 40 tests covering the BaseStore contract
(get / put / update / delete / filter operators / namespace listing / batch /
async / per-namespace eviction isolation), CI on Python 3.11 / 3.12 / 3.13,
and a comparative benchmark in the repo.
Benchmark headline (deterministic synthetic, 12 turns)
| strategy | avg prompt tokens | answer accuracy | warm-hit rate |
|---|---|---|---|
full-history |
52.0 | 0.583 | -– |
vector-only (InMemoryStore + index) |
35.4 | 0.417 | -– |
warm-fallback (WarmStore → vector) |
35.3 | 0.500 | 0.50 |
The synthetic benchmark is best understood as a contract test for the
two-tier pattern. Validating against real agent traces with real embeddings
is the next milestone and is explicitly open work.
What I’m asking from maintainers
Per the contribution guidance, I’d love feedback on:
- Is
BaseStorethe right abstraction? (Versus a Checkpointer or a
Stateextension — I argue it is in the proposal but happy to be
corrected.) - Does per-namespace eviction belong in a community store, or does it
generalize enough that LangGraph would want a built-in policy hook on
BaseStore? - Preferred distribution shape: stay as a third-party package
(warm-memory[langgraph]), or factor out a focused
langgraph-store-warmdistribution mirroring
langgraph-checkpoint-postgres? - Any additional tests or benchmarks you’d want to see before
considering a docs link-out or adoption?
Once the abstraction is settled I’m happy to open a draft PR (docs first,
code second) for whichever path you steer toward.
Thanks for the clear contribution guidance — it made shaping this proposal
much easier.
-– Vivek