
MartinLoop wraps AI coding agents like Claude Code with hard budget caps, verifier gates, and structured audit trails. Each run gets a contract: max spend in USD or tokens, a verifier command like `npm test`, allowed file paths, and a stop condition that isn't just "the agent gave up." The MCP server exposes `martin_run` for governed execution plus read-only tools to inspect run records, attempt history, verifier results, and budget consumption. Reach for this when you need to cap retry loops before they burn through tokens, require proof that tests passed before marking work complete, or build an audit trail showing exactly what changed, what it cost, and why it stopped.
Give coding agents more work. Watch them less. Ship more.
MartinLoop lets you give coding agents real software jobs without babysitting every step.
Your coding agent still writes the code. MartinLoop keeps the job focused, bounded, and accountable — with limits, stop conditions, verification, recovery, and a clear outcome when the run ends.
Trust your coding agents with real work.
Built from thousands of agent runs where the problem was not just intelligence — it was whether you could give the agent a real job and confidently walk away.
Measure your loop tax: npx -y martin-loop@latest audit
Get started: npx -y martin-loop@latest start
Try the demo: npx -y martin-loop@latest demo
MartinLoop is part of the NVIDIA Inception program.
Audit first — run npx -y martin-loop@latest audit to see how much API-equivalent Claude Code spend went into fix-and-retry loops. Add --offline for zero network access or --share for a shareable summary.
Install / onboard — run npx -y martin-loop@latest start, or install it globally with npm install -g martin-loop@latest.
Governed run — define an objective, verifier, budget, and iteration cap with martin run.
Verifier — completion requires fresh verifier evidence bound to the active run and workspace. A configured verifier proves only the checks it runs; VERIFIED is not a claim that the code is bug-free or automatically safe to merge.
Budget — set a hard spend ceiling with --budget-usd and an attempt ceiling with --max-iterations.
Receipts — inspect the latest result with martin dossier --latest and validate stored integrity with martin runs verify --latest.
Hosted sync (optional) — governed work is local-first. Configure MARTIN_API_TOKEN and MARTIN_TELEMETRY_ENDPOINT, then use martin sync status and martin sync flush to send preserved evidence to a dashboard later. See the quickstart.
MCP — install @martinloop/mcp@latest in a supported host or generate host configuration with martin mcp print-config.
Documentation — continue with the quickstart, CLI reference, or MCP setup.
When --model is provided, MartinLoop passes it through unchanged. Without --model, the authenticated host runtime chooses its own default. MartinLoop does not inject a hidden fallback model.
MartinLoop is the system around the coding job.
The coding agent still writes the code. MartinLoop keeps the job, limits, verification, recovery, outcome, and history consistent around the work.
Use MartinLoop when a coding task needs:
Canonical lifecycle:
DEFINE
-> PREFLIGHT
-> CONTROL
-> VERIFY
-> RECOVER
-> PROVE
-> ANALYZE
The product-level flow is Definition of Done -> Controlled Run -> Verified Handoff.
For machine-readable context start with llms.txt, llms-full.txt, and MartinLoop for AI Agents.
Teams should not need to stitch together a separate script or point tool for every part of coding-agent execution. MartinLoop connects the control path around the agent from preflight through post-run evidence.
| Stage | MartinLoop role |
|---|---|
| Define | Capture the objective, verifier, budget, scope, and finish line. |
| Preflight | Check readiness and required workflow evidence before agent spend. |
| Control | Enforce budgets, attempts, path boundaries, policy, and stop conditions while the coding agent works. |
| Verify | Run configured checks and bind the evidence to the active run and workspace. |
| Recover | Preserve recovery and rollback state when another attempt or human review is required. |
| Prove | Produce the authoritative VERIFIED, STOPPED, or NEEDS REVIEW handoff plus receipts. |
| Analyze | Inspect run history, cost provenance, failure classes, dossiers, and shareable evidence after execution. |
MartinLoop does not replace Git, GitHub, CI, dedicated security scanners, observability platforms, code review, or the coding agent itself. It gives those workflows one governed execution record to inspect.
MartinLoop is model-agnostic by design. The job, limits, verification, recovery, and outcome are separate from the worker underneath.
Native coding-agent CLIs and OpenAI-compatible model runtimes use different execution paths, so MartinLoop only describes a worker as fully supported when that execution path has been publicly released and validated.
AI coding agents can create more software than ever. The problem is trusting them with bigger jobs without turning yourself into their full-time manager.
Agents drift. They retry. They break unrelated things. They say "done" too early. And a run that was supposed to save time can create even more work to review.
MartinLoop keeps the job accountable outside the agent itself: clear boundaries, hard limits, configured checks, recovery information, and an explicit outcome.
The goal is simple:
Give the agent more responsibility without giving it unlimited freedom.
More software. Less supervision.
Before changing your workflow, measure it:
npx -y martin-loop@0.7.1 audit
MartinLoop reads Claude Code session history locally and reports how much API-equivalent agent spend happened inside fix-and-retry loops, how often verifier commands failed, the longest retry chain, and sessions that ended red. Session contents stay local. Use --offline to disable even the public pricing lookup, and --share to create a screenshot-ready SVG plus Markdown summary.
Then put limits around the next run:
npx -y martin-loop@0.7.1 run "<task>" --verify "npm test" --budget-usd 2 --max-iterations 3
npx -y martin-loop@latest start
npx -y martin-loop@latest demo
cd martin-loop-demo
npm install
npx -y martin-loop@latest run "Summarize the demo workspace and prove tests still pass" --verify "npm test" --budget-usd 2 --max-iterations 1
Try MartinLoop in a disposable demo workspace:
npx -y martin-loop@latest start
npx -y martin-loop@latest demo
npx -y martin-loop@latest --version
cd martin-loop-demo
npm install
npx -y martin-loop@latest run "Summarize the demo workspace and prove tests still pass" --verify "npm test" --budget-usd 2 --max-iterations 1
npx -y martin-loop@latest dossier --latest
npx -y martin-loop@latest share --latest
Optional global install:
npm install -g martin-loop
martin-loop --version
If this flow is useful, open an issue with feedback so we can keep improving the public experience.
start prints the first-run guided path. run auto-checks doctor, session-start, and preflight, then executes when the environment is ready. Use --proof only when you intentionally want an explicit no-spend lane.
Inspect-first flow:
npx -y martin-loop@latest doctor
npx -y martin-loop@latest session-start
npx -y martin-loop@latest preflight "Summarize the demo workspace and prove tests still pass" --verify "npm test"
share --latest writes three files into the selected run directory under share/: run-receipt.json, run-receipt.md, and proof-card.svg.
Release notes for MartinLoop 0.6.5: MartinLoop 0.6.5.
Release notes for MartinLoop 0.6.6: MartinLoop 0.6.6.
Release notes for MartinLoop 0.6.7: MartinLoop 0.6.7.
Release notes for MartinLoop 0.6.8: MartinLoop 0.6.8.
Release notes for MartinLoop 0.7.0: MartinLoop 0.7.0.
Release notes for MartinLoop 0.7.1: MartinLoop 0.7.1.
MartinLoop governs the job independently of the coding worker.
--engine openai for OpenAI-compatible model endpoints.The worker changes; MartinLoop's budget, scope, verifier, receipt, and integrity contract does not.
More detail: Model and engine support
MartinLoop's terminal presentation is built around the governed lifecycle, not around a single verifier command.
Governed Run Plan shows the configured finish line before work starts, including the task, budget posture, verifier plan, scope, and execution boundaries.
Controlled Run keeps the coding agent working inside those boundaries while MartinLoop tracks attempts, cost, stop conditions, and recovery state.
Verified Handoff closes the loop with one authoritative outcome:
VERIFIED when the configured evidence supports the Definition of DoneSTOPPED when a configured hard boundary ends the runNEEDS REVIEW when completion cannot be established from the available evidenceThe handoff can include verifier steps, scope state, attempt count, cost provenance, unresolved evidence, recovery state, receipt integrity, and the next safe action. The exact fields depend on what the run actually established.
MartinLoop turns an AI coding run into an inspectable execution record: budget used, verifier result, changed files, rollback evidence, and final receipt.
Ungoverned agents can retry until cost and scope drift. MartinLoop adds budget caps, verifier gates, and audit evidence so the run has a clear stop condition.
Long governed runs do not have to mean staring at a spinner. In an interactive terminal, MartinLoop Arcade can be offered while the coding agent continues working in the background.
Arcade is presentation-only. It cannot change the agent, budget, verifier, policy decision, run outcome, or receipt evidence. It stays out of JSON, CI, non-interactive, and other machine-readable execution paths.
Use --arcade to offer Arcade immediately for a supported interactive run, or --no-arcade to suppress it for that run.
Proof receipts are local share bundles for governed AI coding runs. They show the task, spend, budget, verifier result, receipt integrity, and any evidence boundary that should not be rounded into confidence.
This real governed run spent $0.51 against a $3.00 budget. The verifier passed and the receipt integrity was signed, but the proof stayed at EVIDENCE_BOUNDARY because rollback evidence was not recorded.
Generate your own receipt after a governed run:
npx -y martin-loop@latest run "Summarize the demo workspace and prove tests still pass" --proof --verify "npm test"
npx -y martin-loop@latest runs verify --latest
npx -y martin-loop@latest share --latest
Example receipt files: Markdown and JSON.
Use this lane from a clean temp directory to verify the public CLI flow exactly as shipped:
npx -y martin-loop@0.7.1 --version
npx -y martin-loop@0.7.1 start
npx -y martin-loop@0.7.1 demo
cd martin-loop-demo
npm install
npx -y martin-loop@0.7.1 run "Summarize the demo workspace and prove tests still pass" --verify "npm test" --budget-usd 2 --max-iterations 1 --json
npx -y martin-loop@0.7.1 dossier --latest --json
npx -y martin-loop@0.7.1 share --latest --json
For deterministic installs, pin the package line (martin-loop@0.7.1) or use martin-loop@latest. Plain npx martin-loop can resolve a stale local cache on some machines.
Expected share bundle outputs:
share/run-receipt.jsonshare/run-receipt.mdshare/proof-card.svgThe point is not that every governed run is always cheaper. The point is that every run becomes inspectable and enforceable: budget policy, verifier result, stop reason, and evidence are explicit.
For a deterministic public repro lane, use the benchmark workspace and compare governed execution to unbounded retry behavior:
npx martin-loop bench --suite under-3-challengenpx martin-loop bench --suite ralphy-engineering-50A Ralph-style loop is the failure mode where an AI coding agent keeps trying without knowing when continuing is unsafe, uneconomical, or unlikely to succeed.
MartinLoop keeps the useful part of the loop, then adds brakes:
When a coding agent fails, "something went wrong" is not useful enough.
MartinLoop classifies runtime failures into 13 canonical classes so a failed run can tell you what happened, what needs attention, and what future runs can learn from.
See the canonical table: Failure Taxonomy (13 Runtime Classes).
npm test, before a run can count as complete.martin share --latest turns the latest governed run into a local share bundle with a redacted JSON receipt, Markdown recap, and proof-card SVG.| Layer | Purpose |
|---|---|
| Task contract | Objective, verifier plan, repo root, allowed paths, denied paths, acceptance criteria, workspace, project, and budget. |
| Policy and budget | Defaults come from martin.config.yaml; CLI flags can override them. Budget preflight blocks attempts that would exceed policy. |
| Agent adapters | Claude CLI, Codex CLI, Gemini CLI, and direct-provider adapters normalize execution results. |
| Safety and verification | Scope checks, verifier command checks, prompt integrity, and grounding decide whether work can continue. |
| Persistence | JSONL run records, evidence summaries, and repo-backed artifacts make every run inspectable later. Each loop record is locally signed (HMAC, per-runs-root key) and dossier/runs get/runs verify/challenge/badge report an integrity verdict (verified / tamper_detected / unsigned) so post-hoc edits to a record are detectable, not just inspectable. |
actual, calculated, estimated, or unavailable).verified before a run is treated as trustworthy evidence for external review.martin-loop doctor
martin-loop demo
martin-loop session-start [--host <claude|codex|gemini|generic>]
martin-loop phase status|contract|session-start|preflight|run [--execute]
martin-loop preflight <objective> [options]
martin-loop run <objective> [options]
martin-loop bench --suite <suiteId>
martin-loop triage
martin-loop dossier (--latest | --loop-id <id> | --file <path>)
martin-loop runs list|get|attempt|verify ...
martin-loop mcp print-config --host <codex|claude|gemini|cursor|vscode|generic>
martin-loop mcp install --host <codex|claude|gemini|cursor|vscode|generic>
martin-loop mcp verify-install --host <name> [--scope <user|project|local>]
martin-loop mcp rollback --host <name> [--scope <user|project|local>]
martin-loop mcp uninstall --host <name> [--scope <user|project|local>]
martin-loop challenge [--loop-id <id> | --file <path> | --latest]
martin-loop share (--loop-id <id> | --file <path> | --latest) [--out-dir <path>]
martin-loop badge [--format svg|json] [--runs-dir <path>]
Common options:
--budget <n> Hard cost cap in USD
--budget-usd <n> Alias for --budget
--soft-limit-usd <n> Soft budget threshold in USD
--verify <cmd> Verifier command after each attempt
--proof Run verifier-only evidence checks without claiming governed execution
--max-iterations <n> Maximum number of attempts
--max-tokens <n> Maximum token budget
--engine <name> Adapter to use: claude, codex, gemini, or openai
--cwd <path> Repo root for the run
--allow-path <glob> Restrict writes to this path pattern; repeatable
--deny-path <glob> Block this path pattern; repeatable
--runs-dir <path> Override the local Martin runs root
Examples below use npx martin-loop so they work without a global install. If you install martin-loop globally, the martin alias works too.
Use martin-loop share --latest after dossier when you want a redacted bundle you can hand to another person without sending raw run-store files.
More detail: CLI reference and configuration reference.
MartinLoop ships a public deterministic benchmark workspace in benchmarks/ plus the installed-package bench command.
From an installed package:
npx martin-loop bench --suite under-3-challenge
npx martin-loop bench --suite ralphy-engineering-50
From a clean public clone:
pnpm install --frozen-lockfile
pnpm bench:build
pnpm bench:eval
pnpm bench:report:ralphy
Equivalent workspace-filter commands:
pnpm --filter @martin/benchmarks build
pnpm --filter @martin/benchmarks test
pnpm --filter @martin/benchmarks eval
pnpm --filter @martin/benchmarks report:ralphy
The installed-package command reads the shipped public fixtures. The repo-clone workflow runs the public benchmark workspace directly.
Run the standalone MCP package directly:
npx -y @martinloop/mcp
Add it to common hosts:
codex mcp add martin-loop -- npx -y @martinloop/mcp
claude mcp add --transport stdio --scope user martin-loop -- npx -y @martinloop/mcp
claude mcp add --transport stdio --scope user martin-loop -- cmd /c npx -y @martinloop/mcp
Generate host config from the root CLI:
npx martin-loop mcp print-config --host codex --transport stdio --profile minimal
npx martin-loop mcp print-config --host claude --transport stdio --profile diagnostic
npx martin-loop mcp print-config --host gemini --transport stdio --profile full-local
npx martin-loop mcp print-config --host generic --transport stdio --profile github-review
The root martin-loop package, standalone @martinloop/mcp package, plugin metadata, and MCPB product version are aligned at 0.7.1. The MCPB manifest schema remains 0.3.
The public MCP release train labels are:
0.1.4 operator foundation0.2.0 cockpit expansion0.2.5 public MCP package line0.2.7 usability and review release0.3.0 host adoption and onboarding release0.3.1 review and handoff release0.5.3 execution-control and host-compatibility release0.5.5 governed-autonomous execution and proof-surface release0.5.6 hosted run sync, fail-closed rollback, and verified-completion hardeningThe standalone MCP registry/server identifier is io.github.Keesan12/martin-loop.
More detail: MCP setup, MCP tool reference, and MCP compatibility.
npm install martin-loop
import { MartinLoop, createClaudeCliAdapter } from "martin-loop";
const loop = new MartinLoop({
adapter: createClaudeCliAdapter({ workingDirectory: process.cwd() }),
defaults: {
workspaceId: "my-workspace",
projectId: "my-project",
budget: {
maxUsd: 3,
softLimitUsd: 2.25,
maxIterations: 3,
maxTokens: 20_000,
},
},
});
const result = await loop.run({
task: {
title: "Fix auth regression",
objective: "Fix the failing auth regression tests",
verificationPlan: ["pnpm test"],
repoRoot: process.cwd(),
},
});
console.log(result.decision.status);
The root SDK also exports createCodexCliAdapter, createGeminiCliAdapter, createDirectProviderAdapter, and createOpenAiCompatibleAdapter.
More detail: SDK reference and package map.
Requirements:
git clone https://github.com/Keesan12/martin-loop.git
cd martin-loop
pnpm install --frozen-lockfile
pnpm lint
pnpm test
pnpm build
pnpm public:copy-scan
pnpm public:git-surface
pnpm oss:validate
pnpm public:smoke
pnpm release:matrix:local
Standalone MCP validation:
pnpm --filter @martinloop/mcp lint
pnpm --filter @martinloop/mcp test
pnpm --filter @martinloop/mcp build
pnpm --filter @martinloop/mcp smoke:pack
pnpm --filter @martinloop/mcp smoke:published:pack
pnpm --filter @martinloop/mcp verify:release
Issues, bug reports, workflow feedback, and focused pull requests are welcome. Public-facing docs should stay concise, user-centered, and accurate.
git checkout -b feat/your-feature
pnpm lint
pnpm test
git commit -m "feat: describe what you built"
git push -u origin feat/your-feature
Star this repo if you want to give coding agents more real work without babysitting every step.
martinloop.com · support@martinloop.com
MartinLoop is part of the NVIDIA Inception program.
MartinLoop sends minimal anonymous usage data to help improve reliability and prioritize development. On a fresh interactive install, telemetry is enabled by default but remains blocked until the one-time disclosure is shown. After the disclosure, the same run may send the allowlisted anonymous events unless you opt out.
What is sent:
What is never sent:
Endpoint: https://tupopqvqnyyjuxseyxkr.supabase.co/functions/v1/product-events
Headers sent: Content-Type: application/json, User-Agent: MartinLoop-CLI/<version>
No authorization header, API key, or direct table access.
Opt out anytime:
martin telemetry off
Inspect what is sent:
martin telemetry explain
Environment variables that disable telemetry: MARTIN_TELEMETRY_DISABLED=1, DO_NOT_TRACK=1, CI=1
Set MARTIN_TELEMETRY_DEBUG=1 to print the telemetry envelope to stderr without transmitting it.
MartinLoop continues to work normally with telemetry disabled. No features are gated on telemetry consent.
Apache-2.0. See LICENSE.