Recalld is a hosted memory layer for AI agents that extracts atomic facts from conversations or documents, timestamps and reconciles them (updating, superseding, or merging as new information arrives), and returns only the facts relevant to a query. It exposes three API primitives - add, recall, and search - so you can stream turns or documents once and let Recalld handle extraction and reconciliation. “Recall” is a hybrid retrieval plus an LLM curation pass that runs on Recalld’s infrastructure and returns a short, ranked set of passages that directly answer the query; “search” is raw vector similarity without an LLM, for full candidate sets you filter yourself. On the LoCoMo Agent Memory Benchmark, recall returned ~243 tokens at 88.7% accuracy versus search’s 1,627 tokens at 88.2% (6.7× fewer tokens); breakdowns include single-hop 91.7%, temporal 89.7%, open-domain 83.3%, and multi-hop 80.5%.
Operational features emphasize production readiness and compliance: time-aware facts, durable storage with plan-tuned active windows (nothing auto-deleted unless you erase an agent), bring-your-own-key on higher tiers, regional EU/US data residency, TLS/AES-256, and a Data Processing Agreement. Recalld plugs into Model Context Protocol (MCP) clients, offers a Recalld Chat front end that shows which facts were recalled and their cost, and bills via credits (free tier plus Starter $10/mo, Pro $29/mo, Scale $99/mo), with search ~40× cheaper than curated recall. Benchmarks and harness changes are disclosed publicly; exports and erasure are supported.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.