
Gives Claude persistent memory across sessions through a local SQLite store with vector search and cognitive routing. Exposes tools for semantic search over conversation history, session state management, and proactive drift detection that alerts when your agent veers off task. Ships with the prism-coder LLM fleet (1.7B to 32B parameters) for offline tool routing via Ollama, plus optional Brave Search and Gemini integration. Includes a knowledge ingestion pipeline that indexes your codebase into the graph via MCP tools, GitHub webhooks, or REST API. The local tier runs entirely on device with no API keys. Reach for this when you want your AI to remember decisions across weeks without re-explaining context, or when you need tool calling that works offline at zero cost per query.
Give your AI agent memory that lasts — and see the cloud tokens it never had to spend. Persistent sessions, knowledge graphs, offline tool-routing, and an auditable savings meter. Fully local and free.
Prism Coder is an MCP server that gives Claude, Cursor, and other AI tools long-term memory that survives across sessions. It ships with the open-weight prism-coder model fleet (2B–27B) for fast, offline tool-routing — no cloud required. And it keeps score: every call served locally is metered, so prism savings shows the token volume that never reached your cloud model — measured honestly, in tokens.
No account needed. No API keys. Runs on your machine.
A paid subscription adds cloud sync, higher model tiers, and team features through the Synalux portal.
prism savings (or the local_savings
tool from any host) reports the token volume local serving kept off your
cloud model: headline, local share, per-model breakdown. It reports tokens,
never an invented dollar figure, and prints its assumptions and known
undercounts inline — a number you can check, not marketing.route_guard: "local" keeps the
prompt and draft entirely on-device.prism connect configures Claude Code,
Claude Desktop, Cursor, Gemini CLI, and Codex while preserving unrelated
settings.npm install -g prism-mcp-server
prism connect
Use prism connect --dry-run to preview changes, prism connect --all to
configure every detected host, or prism connect --refresh to reconcile
Prism-managed entries after an upgrade. Restart the host after connecting.
Prism works locally without an account, API key, or cloud subscription. Add a Synalux subscription when you want cloud memory, paid-tier skills, or team features.
After a few sessions, ask what it's been worth:
prism savings --period month
💾 Local serving — LAST 30 DAYS
~510K tokens kept off your cloud model
53 call(s) served locally of 58 routed (91%)
Your numbers will differ — that's the point: it reports what your machine
actually served, not a projection. Full report anatomy and the honesty rules
behind it are in the
local_savings section.
Prism also ships as a plugin, which registers the MCP server and the startup skill for you.
Claude Code — from the community marketplace:
/plugin marketplace add anthropics/claude-plugins-community
/plugin install prism-coder@claude-community
Codex — this repository is itself a plugin marketplace:
codex plugin marketplace add dcostenco/prism-coder
codex plugin add prism-coder@prism
The plugin registers prism-mcp via npx -y prism-mcp-server. If you already
configured Prism by hand — prism connect writes an mcp_servers.prism-mcp
entry — you have that server twice under one key. Install the plugin or
run prism connect, not both.
prism connect changes about host subagentsconnect steers bounded work to prism_infer on your machine rather than to
host-spawned agents. What it writes differs per host, and it does not disable
subagents everywhere — Claude Code keeps them and is pointed at an economy
model instead. Prism's local workers stay available over MCP in every case.
| Host | Setting written | Effect |
|---|---|---|
| Claude Code | env.CLAUDE_CODE_SUBAGENT_MODEL = "sonnet" in ~/.claude/settings.json | Subagents stay enabled, pinned to an economy model. Fan-out is discouraged by policy text, not by config |
| Gemini CLI | experimental.enableAgents = false in ~/.gemini/settings.json | Subagents off. Gemini exposes one boolean, so that is all there is to set |
| Codex | features.multi_agent = false in $CODEX_HOME/config.toml (default ~/.codex), plus a bounded fallback: 2 threads, depth 1, cheap subagent model, 900s cap | Subagents off, with a bounded profile underneath so a deliberate re-enable lands somewhere sane |
Two things worth knowing:
experimental is Gemini's namespace, not ours. Prism is not enabling
anything experimental — it writes false to a flag Gemini already defines at
that path. Writing anywhere else would have no effect.enableAgents out of experimental, Prism keeps writing the old path, Gemini
reads the new one, and host subagents quietly turn back on. Nothing errors and
the settings file still looks correct. If you see host subagents running while
enableAgents reads false, check whether the key has moved before assuming
connect failed to write it.Both writes are idempotent in the sense that a host already configured this way
is left untouched — but they are re-applied on every prism connect run,
not only on --refresh. If you deliberately re-enable host subagents, the next
connect will turn them off again. Keep them on by not re-running connect,
or by re-enabling after each run.
prism connect previously installed —
a prism MCP registration in host config no longer counts as consent.
First install of the hook happens through prism connect or not at all.plugins/prism/README.md.prism CLI invocation exit with a commander error at startup (the MCP
server was unaffected). The handoff-sync command is prism handoff …;
prism sync remains cross-backend data synchronization. If you installed
20.17.0, update.prism handoff enable (paid,
off by default), each session_save_handoff seals the handoff to all your
account's device keys and relays the CIPHERTEXT; another machine pulls with
sync_pull_handoff (or prism handoff pull <project>) and opens it locally.prism handoff status|devices to inspect; revoke a lost machine from the portal.local_savings tool + prism savings CLI — the token volume local
serving kept off your cloud model: all time, trailing 30/7 days, or any
--days N window. Tokens, never an invented dollar figure, with the
assumptions and known undercounts printed inline.prism savings --sync-enable uploads per-day
counters only (never content; the payload is a closed field set the server
also enforces); prism savings --team shows the workspace-wide total with
per-member share. Off by default.prism connect --refresh now converges every registration it owns, not
just the top-level one — directory-scoped entries could otherwise keep
launching an old build indefinitely.prism update checks the installed package, not the CLI that happens to
be running, so it can no longer report "current" while the install is stale.PRISM_NO_UPDATE_CHECK=1
opts out.prism autoupdate enable sets up
a daily prism update --if-idle: it updates only the global npm package,
defers while any Prism server is running, and never touches host
configuration — that stays behind a visible prism connect.session_save_ledger/save_handoff calls when its path-to-project
heuristic disagreed with the project you declared — and the registry the
heuristic trusted could contain junk from earlier auto-registration, so
legitimate sessions ended unsaved. Your declaration now always wins; the
disagreement is returned as an advisory warning, and auto-registration
only accepts real repository roots.prism browser captures on macOS were
silently upscaled to the size cap, so a screenshot no longer showed what
actually rendered. Only genuinely oversized captures are resized now, and
the cap no longer clips a standard 1920-wide viewport.prism connect is a converge command. It self-updates first, re-execs,
then reconciles MCP registration, skills, and hooks — no more
"fresh config, stale code" machines.Your skills follow your account. skill_save stores a skill at the
scope you choose: this machine only (local, works offline and signed out),
your account (user — every machine you sign into receives it), or a
workspace (team — shared with members, admin-managed, optionally targeted
to specific people).
Trim the catalog you don't use. skill_manage can release platform
skills you never touch — freeing host skill-catalog budget — and restore
them any time, losslessly. Deleting a scoped skill archives its final
content locally first, so nothing is ever silently unrecoverable.
Delivery that queues instead of failing. Concurrent sessions no longer starve skill sync on the local config store (WAL + busy-timeout) — a failure that previously reported only "partial" where nobody could see it.
Withheld rules still bind. When the context budget can't inline a skill's text, the manifest of withheld names now states that those skills still govern the work and names every way to load them before completion claims.
The budget the floor never spent. A long-standing accounting bug meant no unprotected skill ever inlined at any normal context level — the always-inlined protected floor was debiting the budget meant for everything else. Task-matched skills (like the completion-evidence checklist) now actually arrive.
O_NOFOLLOW descriptor.session_bootstrap
seeds one demo memory and shows it recalled from disk, so the save→recall
loop is felt in session 1. One-shot, contained in its own prism-demo
project, removable with one call.node --checks the built inline script
so an unparseable dashboard can never ship again.npm audit signatures).http:// storage URL is upgraded
to https:// instead of silently sending session content in the clear.prism connect skips
its own registration only when a plugin actually provides prism-mcp
(cache present and enabled), preventing both duplicate and missing
servers.An audit of a real incident (an agent wiped demo data after announcing the wipe — with the ask-first rule committed, bundled, and absent from what any agent actually received) found the protected floor had outgrown every delivery budget: "unprotected" had quietly come to mean "never delivered".
ask-first and feature-preservation join the protected floor (14 → 16).
Protected skills are always inlined; these two now reach every session.--storage accepts auto and synalux — the CLI rejected its own
documented default and the production backend.Memory-grounded answers labelled their sources but never dated them, so a two-year-old note and yesterday's reached the model identically. Nothing in the evidence let it discount the stale one. Prompted by an external review naming the right risk for local-first memory: the data stays local, but bad grounding becomes permanent — storing everything on your machine removes the outside pressure that would otherwise surface a stale note.
Evidence now reads:
[SOURCE 1: ledger:8286581d (recorded 2025-05-29, 431 days ago)]
The date already existed in storage and was being dropped at the snippet layer, so this is plumbing rather than new data collection. Zone-less SQLite timestamps are normalised to UTC — read as local, a ten-minute-old record parsed hours into the future and its age was suppressed entirely, meaning the feature silently did nothing on the freshest memories. An absent or unparseable date renders as nothing rather than defaulting to now; defaulting would make the oldest memories, the ones most likely to be stale, appear freshest.
tests/integration/grounding-staleness.test.ts runs the reviewer's own probe —
seed a deliberately outdated note beside a contradicting fresh one and assert
the model receives both, visibly dated. Anyone can run it.
Not solved, and not claimed: retrieval does not weight recency. A stale note
shown beside a fresh one is the easy case — the model sees both dates and can
weigh them. The hard case is a stale note retrieved alone, because ranking is
by keyword match and an old store returns old results; then the age label is the
only defence and there is no fresher record to compare against. Tracked as
TECH_DEBT.md #4.
Symptom-triggered skills — the rules that fire on "can't see X", "no rows",
"the list is empty" — are meant to load on the turn an incident report arrives.
They never did: every host template called session_bootstrap with {}, so
there was no prompt to match against.
Fixing that raised the question of where matching happens. It now happens
locally. The 28 keyword rules are already public, so there was nothing a local
match could not compute, and callPortal() has no prompt parameter at all —
the guarantee is structural, not a promise. The portal request carries the
project and role only.
A matched rule now arrives as content, not as a name. Native hosts outside the skill-file mirror had no way to read a rule they were only told about, so the rule body is inlined into the startup display, bounded and sized against the real per-project budget.
Setting PRISM_STORAGE=synalux or =supabase with incomplete credentials used
to downgrade silently to local SQLite. The switch was logged to stderr, which
MCP hosts discard, so nothing surfaced it: sessions kept serving stale local
context while the cloud held newer history, and context_source read local
rather than any kind of warning. A session could run that way for weeks.
Naming a backend outright is a strong statement of intent, so it now throws —
naming the missing variables and the PRISM_STORAGE=local opt-out — instead of
quietly splitting your session history. auto is unchanged: it keeps its
documented synalux > supabase > local degradation, pinned by a test.
Upgrade note: if you explicitly set PRISM_STORAGE=synalux|supabase and
your credentials are incomplete, startup now fails with a named error instead
of silently using local data. That error is the fix — set the missing variable,
or choose PRISM_STORAGE=local deliberately. Default (auto) configs are
unaffected.
The throw is deliberately not treated as a recoverable startup fault: that path exists for transient errors (rate limits, 5xx, DNS), which may degrade behind a visible notice. A missing credential is a configuration fault and must not be papered over.
Also: the skill block is now budgeted by default rather than only on request, so a large skill payload cannot crowd out briefing and history.
Security release. Web Scholar scrapes article URLs that come from search-engine output, so the target is attacker-influenceable through SEO poisoning — and because what it scrapes is written into the memory corpus and passed to the configured LLM, a redirection to a local address meant reading an internal service and sending the result onward.
The host guard matched string prefixes instead of parsing the address, and six
spellings of a local address got through: [::1] (URL.hostname keeps the
brackets), 127.0.0.2 (only .1 was enumerated, not all of 127.0.0.0/8),
0.0.0.0, [::ffff:127.0.0.1], localhost. (a trailing dot defeated every
suffix check at once), and [64:ff9b::7f00:1] (NAT64 embeds IPv4 in its low
bits). Host classification now parses addresses and also covers CGNAT,
benchmarking, multicast, reserved, and IPv6 unique-local and link-local ranges.
DNS rebinding is closed too. Every check read the URL string, so a hostname the
attacker controls passed all of them and could still resolve to 127.0.0.1.
Targets are now resolved first, every returned address is validated, and the
connection is pinned to those addresses so the name is never resolved a second
time — which also shuts the window between the check and the connect.
Scrape failures no longer vanish into a bare catch {}, a run is bounded by
PRISM_SCHOLAR_SCRAPE_BUDGET_MS (default 60s) instead of stalling on a raised
article count, and responses are capped at 8 MiB.
This is reachable only when scholar actually runs — scholar_research, or the
background loop under PRISM_SCHOLAR_ENABLED=true — and when the attacker also
controls DNS or a search result. Upgrade if you use Web Scholar.
prism browser could not fail a test. eval 1 === 2 returned status: ok
with exit code 0, a page serving HTTP 500 reported status: ok, and console
errors and uncaught page exceptions were discarded entirely. This release adds
assertions — assert-text, assert-visible, assert-count, assert-url,
assert-title, assert-eval, assert-no-page-errors — that return
status: failed and a non-zero exit. open now reports http_status and
fails on 400 or higher, screenshots are validated rather than assumed, and
eval returns native JSON with its type instead of a Python repr.
The fingerprint layer had never been applied: a wrong keyword argument made
the stealth library throw on every launch — 1,139 failures and 0 successes
since April — while the runner reported it as active. It is fixed, and a layer
that cannot be applied now fails loudly. The headless build no longer
advertises itself through navigator.userAgentData or the Sec-CH-UA header,
and a patch that corrupted Object.getOwnPropertyDescriptor on every page
under test has been removed. These remain best-effort test aids, not a
guarantee against bot detection.
--local-only now actually isolates: WebSocket, EventSource, WebRTC and
sendBeacon egress bypass request routing and were never blocked, and service
workers were allowed through. --cleanup was a no-op in the two modes agents
use. Site isolation, phishing detection and popup blocking are no longer
disabled by default, since these profiles hold live authenticated cookies.
New for test runs: --ephemeral-profile and --storage-state for hermetic
authenticated flows, pages/switch-page so OAuth popups are reachable,
--fail-fast, --fast, --trace/--video/--har, and
profiles --prune-older-than for profile maintenance.
Prism now remembers that a conversation successfully loaded its project context
when the MCP server restarts or another Prism process handles the next request.
session_save_ledger and session_save_handoff no longer fail with a false
context_not_loaded error in that flow.
The recovery remains fail-closed: authorization is limited to the exact project and conversation, expires with the existing context window, and stores no plaintext conversation identifier. Cross-project, forged, malformed, expired, or future-dated receipts are still rejected. The release also updates PostCSS to the patched 8.5.23 release.
session_search_memory on the portal tier (Synalux-backed installs) now
fuses semantic similarity with exact-term lexical matching via weighted
reciprocal-rank fusion. Measured on blind probes against a real
8.5k-entry corpus: fused retrieval was never worse than semantic
alone at top-5, and exact identifiers — TPNs, function names, error
strings — now rescue queries that embedding similarity blurs. Results say how they were found — hybrid retrieval
headers, per-hit sem#/lex# arms — and a lexical-only rescue is labelled
exact-term match instead of pretending to a similarity score. Local
SQLite installs keep pure vector search; hybrid needs the portal's
lexical index.
prism connect now reads Claude, Cursor, Gemini, and Codex configuration
through a single verified file snapshot, preventing another process from
swapping a file between Prism's safety check and its read. Supported symlinked
dotfiles still work, while dangling or planted symlinks fail loudly instead of
being followed or overwritten. This release also carries the patched
dependencies and cross-platform release checks introduced in v20.2.5.
Cloud fallback is now documented consistently as Gemini 3.6 Flash. Plan
ceilings govern automatic prism_infer routing; direct use of any downloaded
model through local Ollama remains free on every tier.
Greeting-only assistant replies are skipped before ledger writes. Existing
greeting rows are filtered at read time across native startup, MCP context, and
prism load --json, while entries containing decisions, TODOs, changed files,
or non-session events remain visible. Historical rows are not destructively
deleted. If Synalux has a transient startup failure, Prism displays one bounded
local last-good snapshot and clearly labels it; permanent authorization or
validation failures still fail loud, and later writes remain cloud-routed.
prism connect now installs one orchestration contract for Claude Code,
Claude Desktop, Cursor, Gemini CLI, and Codex. Bounded delegated work goes to
session_task_route and the local prism_infer worker first; routine work must
not create background host agents. Local workers can receive the active
project's dashboard-configured quick, standard, or deep memory and select a
RAM-safe 2B/4B/9B/27B model at call time. The router forwards complexity but
does not choose the model; prism_infer owns the final decision using memory
and context fit, installed models, live RAM, entitlements, and explicit caller
overrides.
Codex and Gemini native agent fan-out are disabled during connect. Codex keeps
a two-thread, one-level Terra/low fallback profile if the developer explicitly
re-enables native agents later. Claude Code keeps native agents as a last-resort
path but pins their model to Sonnet. Cursor and Claude Desktop do not expose a
supported global subagent-policy file, so they receive the identical workflow
through Prism's MCP server instructions. prism_infer safety boundaries and
the host's final verification responsibility are unchanged.
prism connect now downloads the authoritative Synalux skill manifest and
materializes entitled packages in the native ~/.agents/skills directory
before the command exits. Codex therefore sees the current skillset on its
first launch instead of requiring a second restart. Prism rechecks the same
snapshot at MCP startup, session load, and every five minutes—without host
lifecycle hooks.
On the first user turn, Prism's native skill, MCP metadata, and managed host
instructions request one session_bootstrap({}) call. Prism then uses the
dashboard's developer name, Auto-Load Projects, and quick, standard, or deep
setting. The response stays focused on greeting and session state because tier
skills are already present in the host's native skill directory.
Hook-free MCP can provide and prioritize that ready-to-display block, but the host model still owns the final assistant message and may summarize it. Prism does not claim a deterministic verbatim greeting on third-party chat surfaces; that would require a host lifecycle hook, launcher, extension, or Prism-owned panel. Context loading itself remains complete even when a host shortens the visible reply.
Free accounts receive only the public hook-free prism-startup package; the MCP
server still supplies a compact, non-proprietary safety and evidence contract.
Authenticated paid accounts receive the protected behavioral and engineering
packages plus the current subscribed routing set. The paid
evidence-first-protocol keeps ordinary coding lightweight: one correlated
reproduction is enough to begin an edit, while strict acceptance starts only
before a completion claim, push, or release and inspects only the exact artifacts
used as proof. Upgrades install newly entitled packages; verified downgrades
remove only Prism-owned packages while preserving local skills and locally
modified conflicts.
When upgrading an older Claude Code installation, prism connect removes only
the exact Prism-owned startup, skill-sync, handoff, and drift hook actions from
the legacy bootstrap. It also removes the recognized legacy Prism startup
sections from ~/CLAUDE.md, preserves every other instruction, and installs a
small ownership-marked native block that selects session_bootstrap({}) on the
first turn. User hooks, custom instruction sections, and near matches remain
untouched; native skills and server-side reminders preserve those Prism
features without host lifecycle hooks. Because hosts expose no native
session-end callback, handoff at shutdown is instruction-driven rather than a
guaranteed lifecycle event.
After Claude Code's native user registration succeeds, the same default or
--refresh command checks the nearest .mcp.json from the current directory
through the home directory. It removes only the exact legacy
prism-mcp entry { "command": "npx", "args": ["-y", "prism-mcp-server"] }
that would otherwise shadow the user registration. Custom Prism entries and
their additional fields, plus unrelated servers, are preserved; malformed
files fail loud without changes. --dry-run reports the recognized migration
without changing the file.
prism connect now carries an explicit PRISM_STORAGE=auto|local|synalux|supabase
into every managed host registration and rejects invalid values before changing a
config file. In auto, a portal-confirmed free tier uses local SQLite, while
Standard, Advanced, and Enterprise use Synalux cloud memory. If entitlement
resolution is unavailable, Prism fails closed instead of splitting history across
backends. Storage remains independent of local-first model routing.
Install Prism globally and run prism connect. It detects Claude Code, Claude
Desktop on macOS, Windows, and Linux (beta), Cursor, Gemini CLI, and Codex, then safely registers the
server from the installed package. Existing custom entries are untouched;
--dry-run previews changes and --refresh updates only Prism-managed entries.
prism_infer gains a failure contract: pass escalation: "report" and every call returns a structured gate_outcome — success, degraded (gate-failed output served anyway, explicitly flagged), or refused (typed, with reason, instead of a thrown error). Degraded output can no longer serve silently.
Prompts over 4000 chars were blanket-refused when cloud was off. Now the full text gets a deterministic reserved-keyword scan plus a head+middle+tail excerpt classification — clean oversize prompts serve locally with a distinct UNCERTAIN_LENGTH audit marker. Clinical/reserved handling is unchanged (and its keyword floor got stronger).
Tier context limits now match the live Modelfiles (27b/9b are 4096-token models; 4b/2b are 32768 — the old table had it backwards). Tiers that can't hold your prompt are skipped with a visible ctx_insufficient reason; if nothing fits, you get the full prompt on cloud or a loud error — never an answer computed from a silently-clipped prompt.
Entitlements carry a source field: portal (real), unconfigured (free by design), or fallback_free (portal unreachable — free limits ASSUMED). Pass strict_entitlements: true to fail loud instead of running degraded.
The verify_behavior tool crashed on every call (-32602 expected object, received string) — the handler returned a bare string instead of an MCP CallToolResult object. Fixed, with contract + fail-closed regression tests so the safety gate can never silently break again. If you're on 20.0.6/20.0.7, update.
Reserved clinical content is now Claude-or-refuse (never served by a smaller model than the one that refused it), skill delivery gained a JWT auth fallback (paid-tier skills now reach machines using only PRISM_SYNALUX_API_KEY), and every prism_infer call is recorded in a persistent infer_metrics ledger. Full details in CHANGELOG.md.
The local-inference-first skill covers 15 hard-trigger categories (code gen, regex, format conversion, summarization, documentation, factual lookup, classification, shell commands, config gen, and more). Pasted code blocks now trigger delegation regardless of question phrasing. Measured delegation rate: 30-35% on engineering sessions, 40-60% on transform/content sessions. Rate depends on prompt mix, not the skill — the instruments now self-validate with nonDelegatedCount to prevent curated-set tautologies.
Qwen 3.5 models (9B/27B) with thinking enabled could burn all tokens on <think> blocks and return empty content, causing a cascade to 4B. Now detects think-only responses and retries the same tier with thinking disabled — preserving model quality instead of falling to a smaller model.
The reserved-category classifier now retries once with a longer timeout on cold-model failure, then falls back to a deterministic keyword backstop before refusing. Over-length prompts (>4K chars) are classified as UNCERTAIN before reaching the classifier — prompt padding can no longer force the ERROR branch. This eliminates the cold-start refusal problem without weakening the safety gate.
When the LLM classifier fails (timeout, injection, resource pressure), a deterministic regex floor catches reserved vocabulary (restraint, seclusion, self-harm, suicide, overdose, crisis de-escalation, etc.) including inflected and verb forms. Blocks prompt-padding and classifier-injection attacks on the ERROR path.
The safety statement in the MCP server instructions field now imports from boundaries.ts — one source of truth instead of two hand-maintained copies. Boundaries version bumped to v3 with an explicit delivery decision documented in code.
The Layer 1 semantic classifier now runs for every user, not just paid tiers. Reserved clinical content is refused on free tier when cloud is unavailable — fail-closed.
session_save_ledger deduplicates identical entries within a 5-minute window.
scripts/generate-evidence.sh regenerates all 5 evidence files with built-in assertions. Run bash scripts/generate-evidence.sh to verify the full pipeline.
Prism MCP is now Apache-2.0. The thin-client architecture means all proprietary value (skill resolution, tier gating, billing, cloud inference) lives server-side — the open client carries no moat to protect. Apache-2.0 removes the enterprise adoption friction that AGPL caused.
Skill routing, budget management, and content resolution have moved server-side to the Synalux portal. The MCP client is now a thin API caller — simpler, smaller, and portable across any host (Claude Code, Gemini, Cursor, autonomous scripts). Offline fallback reads the last successful response from local SQLite.
The Voyage AI embedding adapter was independently reimplemented from the Voyage API docs to ensure 100% project-owned copyright. Default model updated to voyage-3.5. See PROVENANCE.md for details.
Session drift detection (GATE 5) no longer requires Claude Code hooks. The timer runs server-side per conversation, piggybacked on every MCP tool response. Works for any host.
External contributions now require signing the Individual CLA. The CLA check is merge-blocking on the main branch.
The free tier needs no account, no API key, and no cloud. Install Prism, then register it with every supported MCP host already installed on your machine:
npm install --global prism-mcp-server
prism connect
prism connect detects Claude Code, Claude Desktop (macOS/Windows/Linux), Cursor,
Gemini CLI, and Codex.
Use prism connect --all to target all five, --host <name> for one host, or
--dry-run to preview the files that would change. Existing prism and
prism-mcp entries are never overwritten by default. --refresh updates only
an entry previously created by Prism; custom entries remain untouched.
For Claude Code, both the default command and --refresh also remove the exact
legacy project-scoped npx -y prism-mcp-server entry from the effective
ancestor .mcp.json after the native user registration succeeds. No custom or
near-match project entry is changed.
Close the target MCP hosts before a non-dry-run registration so they cannot
edit their configuration at the same time.
The same connection installs the local-first orchestration contract:
| Host | Managed containment |
|---|---|
| Codex | features.multi_agent=false; a 2-thread, depth-1 Terra/low fallback profile is retained for explicit re-enable |
| Gemini CLI | experimental.enableAgents=false |
| Claude Code | CLAUDE_CODE_SUBAGENT_MODEL=sonnet; managed instructions reserve it for last-resort fallback |
| Cursor | Canonical policy delivered through MCP initialize instructions |
| Claude Desktop | Canonical policy delivered through MCP initialize instructions |
All five receive PRISM_AGENT_POLICY=local-first in their managed Prism MCP
entry. Routine tasks use the RAM-aware local worker; native/background fan-out
is not the default workflow. session_task_route supplies a complexity hint;
prism_infer remains the single owner of model and thinking selection and can
choose 27B when its viability gates support it.
Set PRISM_STORAGE before running prism connect to preserve an explicit
storage choice in the generated host entries. This does not change local-model
routing; Synalux cloud storage separately requires an active cloud-memory
entitlement.
Codex registration preserves unrelated ~/.codex/config.toml content, appends
only the marked Prism MCP block, and updates only the documented local-first
feature/agent keys. CODEX_HOME is respected when set and must already exist,
matching Codex's own contract. Restart Codex CLI, the
IDE extension, or the ChatGPT desktop app after connecting.
Restart the connected host and your agent now has memory backed by a local
SQLite database (~/.prism-mcp/data.db). See IDE setup
for manual configuration and host-specific paths.
Optional — local model fleet for offline tool-routing. Pull whichever fits your hardware:
ollama pull dcostenco/prism-coder:2b # 3.3 GB · on-device / lowest RAM · sees images (100% on our routing suite)
ollama pull dcostenco/prism-coder:4b # 3.5 GB · verifier · sees images (100%)
ollama pull dcostenco/prism-coder:9b # 6.7 GB · default router · sees images (95.7%, reasons before answering)
ollama pull dcostenco/prism-coder:27b # 16.8 GB · complex code / quality · text only (100%)
Prism detects both the namespaced (dcostenco/prism-coder:9b) and bare (prism-coder:9b) Ollama tags automatically.
The 2b/4b/9b tiers carry a vision tower and accept screenshots through
prism_infer({ images: [...] }) — pass absolute paths or base64. Image
requests are refused rather than answered blind when no tier (or the Layer 1
safety classifier) can actually see the image, so a text-only model is never
handed a prompt about a screenshot it never received. The 27b is text only.
Your AI agent forgets everything between sessions. Prism fixes that — and adds verification, drift detection, and multi-agent coordination on top.
Every conversation feeds a persistent store. The next session loads the right context automatically — no re-explaining.
The dashboard shows your current project state, pending TODOs, intent health, and a neural knowledge graph — all built automatically from your agent sessions.
It runs on loopback and is gated by a per-startup token by default — open the
tokenized URL printed in the startup log (http://localhost:3000/?token=…).
Requests with an untrusted Host/Origin are refused, closing the DNS-rebinding
exposure fixed in GHSA-9cvx-7x8q-3g6m. See docs/IDE_SETUP.md
to pin the token, disable it, or configure Basic Auth / JWKS.
session_export_memory writes your memory out as plain files you can read,
diff, and commit. Nothing goes through a model to produce it.
markdown human-readable — drop it in a PR to show what the agent actually did
json machine-readable — import into another Prism instance
vault zipped Markdown with YAML frontmatter and [[wikilinks]] (Obsidian, Logseq)
This is the surface to reach for when you want to answer "did the agent verify this, or is it claiming it did?" — the export is a record you review after the fact, in a diff or a pull request, rather than a live view you have to go and open. The same data is available from the dashboard's Export ZIP and Export Vault buttons.
Ask "what did I decide about the auth flow last month?" and get an answer with citations, combining vector similarity, full-text search, and graph traversal.
Every session is logged with files changed, decisions made, and TODOs. Search, filter, and replay any past session.
Every prism_infer call tracks which model handled it (local Ollama vs cloud) and how many tokens were consumed. When you save a session, Prism shows a summary:
📊 Inference Metrics (this session):
Total calls: 12 — Local: 10 (83%) | Cloud: 2 (17%)
Prompt tokens: 7,840 evaluated / 8,420 submitted est.
Completion tokens: 3,150
Cloud tokens saved (est.): 11,570 — token volume handled locally instead of cloud
Avg latency: 1,240ms
By model:
prism-coder:27b: 6 calls, 7,200 tokens, avg 1,800ms
prism-coder:9b: 4 calls, 2,870 tokens, avg 620ms
synalux-27b: 2 calls, 1,500 tokens, avg 1,100ms
Cloud tokens saved is the honest routing metric — it accrues only when local Ollama handles a call that would otherwise have gone to Synalux cloud inference. A compact version appears inline after every 5th prism_infer call: 📊 local 10 (83%) · cloud 2 (17%) · ~11,570 tok · avg 1,240ms · 11,570 cloud tok saved.
Local calls use actual Ollama token counts (prompt_eval_count / eval_count from Ollama); cloud calls use char/4 estimates. Metrics are tracked locally — no portal dependency, no env vars, works offline. Per-call data is also forwarded to the Synalux portal as best-effort analytics (independent of the display).
Long agent sessions can wander from their original goal. session_detect_drift compares current work against the stated goal and returns on_track / minor_drift / major_drift so the agent can self-correct.
AI agents apply patterns from checklists without understanding the real-world impact. The verify_behavior tool challenges the agent with a scenario it must answer before editing — forcing it to think through what the end user will experience.
Agent: "I'll revert this kitchen display change"
Prism: "⚠️ Scenario: A cook sees a 3-item ticket. One item is voided.
What should the cook see after the void?"
Agent: "The ticket stays visible with the remaining 2 items."
Prism: "Correct — your revert would hide the ticket entirely."
17 built-in domains (billing, auth, ordering, clinical, HR, and more). Custom domains per workspace on Enterprise. No hooks needed — works in any MCP client.
Roll back to any previous session state. Compare diffs between versions. Restore a known-good state with one click.
Three memory types, automatically sorted: episodic (what happened — session logs, decisions), semantic (what's true — facts, architecture), and procedural (how to do X — workflows, patterns). When you search, the router picks the right store instead of dumping everything.
Coordinate multiple AI agents working on the same project. Each agent has its own session, but they share memory through the knowledge graph. The Hivemind Radar shows real-time agent status, tasks, and activity.
Search across all memories with highlighted results, knowledge graph editing, and memory density metrics.
The free tier runs entirely on your machine. Paid tiers add cloud sync through the Synalux portal, which is what enables cross-device memory and team sharing.
| Local tier (free) | Cloud tier (paid) | |
|---|---|---|
| Memory storage | Local SQLite | Synalux portal (Supabase-backed) |
| Inference | Local Ollama models | Local models + Gemini 3.6 Flash fallback |
| API keys required | None | Synalux subscription key |
| Web search / scrape | Not included | Via Synalux portal (provider keys server-side) |
| What leaves your machine | Nothing | Memory text, file paths, search queries, and inference prompts/drafts when their cloud feature is used, sent to the portal over TLS. Cloud memory writes are PHI-redacted; inference and route requests are transient. |
| Works offline | ✅ | Local features yes; sync/cloud no |
Handling sensitive data. Cloud memory writes pass through automatic
redaction (SSNs, dates of birth, medical record numbers, phone numbers, emails,
and clinical identifiers are stripped before storage). Cloud inference and
route correction send the request over TLS for processing and do not store it
as Prism memory; use route_guard: "local" or the local tier for a full
air-gap. Enterprise includes a HIPAA Business Associate Agreement.
The prism-coder fleet uses Qwen3.5 for MCP tool-routing AND general inference. The 9B and 27B are fine-tuned; the 2B and 4B use stock Qwen3.5-4B at different quantization levels. The 27B scored 100% on our internal 115-case tool-routing suite and 100% on an internal 15-problem coding eval, at $0 inference cost. These are self-run evaluations, not BFCL leaderboard submissions.
prism_infer supports three modes: route (tool routing, fast), chat (conversation) and code (code generation). Reasoning is decided by the tier, not the mode: a tier carrying MODEL_TIERS.prefersThinking also carries a minLocalTokens floor so reasoning cannot crowd out the answer, and only those tiers use <think> blocks (stripped before the response is served). The 9B does; the 4B and 2B do not, because on those tiers reasoning drew down the same num_predict budget the answer needed and returned an empty response. An explicit think: true still overrides, for a caller who has sized max_tokens for it. If the local model fails a quality gate (empty, think-only, or truncated), paid tiers automatically escalate to Gemini 3.6 Flash via the Synalux portal.
Every route-mode result is parsed locally and checked against allowed_tools
before it reaches the host. Malformed or unadvertised calls become NO_TOOL.
With route_guard: "auto" (the default), Standard and higher plans also send
a well-formed draft for one of Prism's seven trained tools—or an unadvertised
draft that may need correction—to Synalux for authenticated deterministic
correction. Advertised custom host tools remain local. Set
route_guard: "local" for a fully on-device route path.
| Model | Ollama tag | Size | Vision | Routing accuracy¹ | Role | Automatic routing tier |
|---|---|---|---|---|---|---|
| Qwen3.5-4B Q4_K_S | prism-coder:2b | 3.3 GB | ✅ | 100% | On-device / lowest RAM (4.5 GiB free) | Free |
| Qwen3.5-4B Q4_K_M | prism-coder:4b | 3.5 GB | ✅ | 100% | Verifier (5.2 GiB free) | Free |
| Qwen3.5-9B (LoRA) | prism-coder:9b | 6.7 GB | ✅ | 95.7%² | Default router / workhorse (9 GiB free) | Standard+ |
| Qwen3.5-27B (LoRA) | prism-coder:27b | 16.8 GB | — | 100% | Complex code / quality (21 GiB free) | Advanced+ |
¹ Self-run on a narrow 115-case MCP tool-selection suite, temperature: 0,
measured through the call path prism_infer actually uses (/api/chat, each
model's own template). It says these models pick the right tool on our own eval,
nothing more — not a general capability measure, and not an independent
benchmark result. Earlier revisions of this table quoted 99.1–100% from a
harness that hand-rolled a ChatML prompt with raw: true, bypassing the
template; those numbers described a path no caller exercises. Full methodology
caveats below.
² The 9B is the one tier that reasons before answering, and it is measured with
reasoning enabled: 95.7% with thinking, 83.5% without. prism_infer sets this
per-tier (MODEL_TIERS.prefersThinking), so callers get the 95.7% path by
default. Reasoning costs roughly 600 tokens, which is why the 9B also carries a
2,048-token local floor.
Vision. The 2B/4B/9B tags ship a separate projector layer (0.68–0.92 GB)
and read images; the 27B is text-only. prism_infer probes for that layer and
skips a tier with no vision rather than sending it an image — asked directly, a
text-only model will still answer confidently about pixels it never received.
Exercised against the real models in tests/integration/visionScreenshot.test.ts.
These tiers control automatic prism_infer selection, not Ollama itself. Any
user can run any downloaded on-device model directly through Ollama on every
plan.
Weights: huggingface.co/dcostenco (public GGUF). Latency depends on model size and hardware — see Benchmarks to measure it on your own machine rather than trusting a printed number.
query → prism-coder:9b (local router, default)
→ prism-coder:4b (grounding verifier)
→ prism-coder:2b (iPhone / mobile, auto-selected by RAM)
→ prism-coder:27b (complex tasks, on demand)
→ Gemini 3.6 Flash cloud fallback (paid tiers, for max quality)
Route output and evidence-grounded answers use separate gates. Every tier gets the local route parser and advertised-tool registry; Standard and higher plans can add the private deterministic route correction. Evidence verification is opt-in (or automatic when evidence is supplied) and remains separate from route selection.
| Layer | What | Model | Cost |
|---|---|---|---|
| L1 | Crisis/medical safety gate | None (regex) | 0 ms |
| L3-Registry | Envelope validation + advertised-tool enforcement (all tiers) | None | 0 ms |
| L3-Route | Authenticated deterministic route correction (Standard+) | None | Network latency |
| L3-Tier0 | Integer grounding (set membership) | None (deterministic) | 0 ms |
| L3-Tier2 | NLI verifier (claim → ENTAILED/NEUTRAL/CONTRADICTED) | prism-coder:2b | ~200 ms |
| L4 | Hallucination judge (opt-out for clinical) | prism-coder:4b | ~500 ms |
Fail-closed on the verified path: when the grounding verifier runs, timeout, ambiguity, or missing evidence yields a refusal, not pass-through. If the paid route correction is unavailable, the local registry still blocks malformed and unadvertised calls and reports an allowed preserved route as degraded.
Published benchmark numbers are concise summaries of internal deterministic evaluation. Evaluators, exhaustive cases, exact tier-routing matrices, and raw model outputs stay in the private engineering repository and are not included in the npm package or public source tree.
Routing evaluation. On a narrow tool-selection suite, the fleet achieved near-saturated results across three seeds. This measures offline MCP routing reliability, not general model capability.
| Model | Routing accuracy | Notes |
|---|---|---|
| prism-coder:2b (Q4_K_S) | 100% | The 2B was requantised when vision shipped; the old 99.1% was Q3_K_M |
| prism-coder:4b | 100% | |
| prism-coder:9b | 95.7% with reasoning | 83.5% without — the only tier where this differs |
| prism-coder:27b | 100% | |
| Claude (frontier, same eval) | ~98% | Stronger everywhere outside this narrow task |
Measured through /api/chat with each model's own template — the path
prism_infer uses. temperature: 0, so the three seeds only reshuffle case
order and cannot disagree; earlier revisions cited that agreement as
confirmation, which it never was.
Memory uplift (LoCoMo-Plus, self-published). A separate long-context dialogue benchmark (dcostenco/Locomo-Plus) measures how much structured memory helps a base model retain multi-day context. Results show large gains when a model is paired with Prism memory versus running raw. Note this benchmark is authored, run, and LLM-judged by this project — treat it as a reproducible demonstration, not an independent third-party result, and run it yourself with the commands in that repo.
Code generation evaluation. In a small July 2026 deterministic execution check, the local 9B passed 2/3 tasks; the local 27B and Gemini 3.6 Flash each passed 3/3. This is a self-published regression signal, not an independent leaderboard or a claim of broad model equivalence.
cloud_fallback: true)Prism always tries an eligible local model first. If the quality gate detects an empty, truncated, think-only, or looping response, paid tiers can retry the request through Gemini 3.6 Flash. Free-tier routing stays local and reports the quality-gate outcome without making a cloud call.
Product capabilities and plans change frequently. The comparison below is intentionally limited to publicly documented differences; it is not a claim that another product lacks an unlisted feature.
Legend: ✅ documented, ◐ conditional or plan-dependent, — not compared, ? verify with the provider.
BRAVE_API_KEYsecretBrave Search API key for web search tools
GEMINI_API_KEYsecretGoogle Gemini API key for AI-powered summarization and embeddings
SUPABASE_URLSupabase project URL for session memory and knowledge persistence
SUPABASE_KEYsecretSupabase service role key for database access
PRISM_USER_IDUser ID for multi-tenant row-level security isolation (defaults to 'default')
GCP_PROJECT_IDGoogle Cloud project ID for Vertex AI Discovery Engine integration
VERTEX_LOCATIONVertex AI location/region (defaults to 'global')
VERTEX_DATA_STORE_IDVertex AI Discovery Engine data store ID for enterprise search