Esment logo
Esment
Esment Enterprise

Memory you own, run entirely inside your own walls

Deployed on your infrastructure, with the option to go fully offline. Local LLM, embeddings, and storage — zero outbound calls.

Request a demo

Pricing is scoped to your deployment.

On-prem or air-gapped

Your hardware, your rules

One signed binary

LLM + embedder + storage

Full audit trail

Every mutation, a commit

No lock-in

SQLite you can open yourself

esment-enterprise.dmgsigned ✓
esment (server binary)
qwen2.5 · GGUF, bundled
fastembed-rs · local vectors
SQLite · knowledge graph

0 outbound connections required

Nothing leaves.
Unless you say so.

Most “self-hosted AI” still phones home for the model. Esment doesn’t — inference, embeddings and storage all run on your hardware by default. Every default is a choice: swap in a cloud model for one workload, keep another fully air-gapped.

InferenceBundled llama-server + GGUFOr your OpenAI / Ollama endpoint
Embeddingsfastembed-rs, computed on-boxOr any API you already pay for
StorageSQLite on your diskOr Postgres in your VPC
TelemetryOff. There is nothing to turn off.

Your institutional memory is strategic. Don’t leave it in someone else’s cloud.

esment audit --followlive
c41f9a2updateretention policy — 90d → 30d
8e02d7bconfirmreaffirmed by second source
3ba61c9supersedeolder contradicting fact archived
f19c04ecreateextracted from #contracts thread
a70b3d5readtime_travel --at 2026-05-01
every mutation is a SHA-256 commit — exposed over REST, MCP, and both SDKs

Every action, on the record.

Role-based keys

Admin, read/write, or read-only, enforced by default. Read-only keys get a hard 403 on any mutation.

Tenant isolation

Every row is scoped by tenant. One team’s key can’t see another’s memories — by construction.

Bi-temporal history

Nothing is silently overwritten. Diff any two states, or ask what the system believed at any point in time.

Inspectable, not a black box.

A dedicated desktop app ships with every deployment — browse the knowledge graph, edit memory, and walk the audit history. No SQL console required.

Esment enterprise app

Galaxy view

Every memory and its relations, explorable in 2D

Memory blocks

Human-editable evergreen context, versioned

Audit timeline

History, diffs and time-travel — visual

Scale

One deployment.
Every team walled off.

Multi-tenancy is first-class. Every tenant gets its own keys, quota, and slice of the knowledge graph — split into four kinds of space, so an org-wide fact and a private preference never collide.

tenant · product-teamkeys · quota · graph
orgdecisions, policies, client preferences — shared
personaleach person’s own context, private to them
projectscoped per client or workstream
agentmemory for autonomous workflows
tenant · legalkeys · quota · graph
isolatedown hashed keys, own quota, own slice of the graph

People leave.
The context stays.

Decisions, clients and reasoning live in your organisation’s memory — not in individual chat histories that walk out the door with the person who wrote them.

Onboarding starts from everything the team already knows. Day one, not month three.

How memory is shared

Fits what you already run.

One memory store, every surface your team could want. Point an existing integration at the proxy and memory gets injected — no code changes.

MCP · stdioClaude Desktop, CursorMCP · HTTPhosted assistantsOpenAI-compatible proxyzero code changesREST APIanything elsePython SDKtyped clientTypeScript SDKzero depsOAuth 2.1PKCE + DCR + JWKSWebhooksSlack, CRMs, warehouses

Support & SLA

A direct line, not a queue.

A contract, an SLA, and a direct channel to the people who maintain the retrieval pipeline — not a support script. We help plan the rollout and stay reachable after it ships.

  • Everything in Teams
  • Observational memory — saves passively from the conversation stream
  • Conversation ingestion — POST raw turns, extraction runs asynchronously
  • Webhook events — Slack, CRMs, data warehouses
  • Dedicated signed binary, built for your deployment
  • Contract + SLA on response times and uptime

Faster than a blink.

A blink takes ~150 ms

Recall is deterministic code, not another model call — so memory can sit inside every single request without you ever noticing it’s there.

Full-text matchstage 1 · exact words
< 5 ms
Semantic searchstage 2 · similar meaning
< 20 ms
Graph expansionstage 3 · connected memories
< 30 ms
Full recall, end to endall four stages, reranked
< 50 ms

Bring your own infrastructure.
We’ll help you deploy.

Cloud, on-prem, or air-gapped — tell us about your environment and we’ll scope it together.