CCM
/MCP
SkillsMCPMarketplacesDigestToolsAdvertise

This week in Claude

Every Monday: Claude Code, Agent SDK, MCP, and the Anthropic platform moves worth your time.

Skills by Category
Frontend DevelopmentBackend & APIsTesting & QASecurityDevOps & CI/CDGit & Pull RequestsDocumentationCode Review & QualityAI & Agent BuildingSkill Development
MCP Servers by Category
Sales & MarketingWeb & Browser AutomationDatabasesAI & LLM ToolsCloud & InfrastructureCommunication & MessagingDeveloper ToolsDesign & CreativeDocuments & KnowledgeSearch & Web Crawling
Marketplaces by Category
AI Agents & OrchestrationLLM IntegrationDevelopment ToolsFrontend & UIBackend & APIsDatabasesTesting & Code QualityDevOps & CloudSecurity & ComplianceGit & Version Control

Claude Code Marketplaces

Discover Claude Code plugins, extensions, and tools. Automatically updated directory of Anthropic Claude AI marketplaces with development tools, productivity plugins, and integrations.

Resources

  • Browse Skills
  • Browse MCP Servers
  • Browse Marketplaces
  • Skill index
  • MCP index
  • Marketplace index
  • Plugins Reference

Community

  • About
  • Tools
  • Feedback
  • Privacy Policy
  • Advertise

Built for the Claude Code community with Claude Code by mertbuilds.com

Independent project, not affiliated with Anthropic
comet-ml avatar

Opik MCP Server

comet-ml/opik-mcp
20810 toolsauthSTDIOregistry active
Summary

Opik Mcp provides an MCP server that integrates Opik, a prompt engineering and LLM observability platform, with MCP-compatible clients through both local stdio and remote streamable-http transports. The server exposes tools for prompt lifecycle management, workspace and project exploration, trace analysis, metrics and dataset operations, and MCP resources for metadata-aware workflows. It solves the problem of enabling AI assistants and development tools to programmatically interact with Opik's LLM evaluation and monitoring capabilities without requiring direct API integration.

Install to Claude Code

verified
claude mcp add opik-mcp --env OPIK_API_KEY=YOUR_OPIK_API_KEY --env OPIK_API_BASE_URL=YOUR_OPIK_API_BASE_URL --env OPIK_WORKSPACE_NAME=YOUR_OPIK_WORKSPACE_NAME -- npx -y opik-mcp

Run in your terminal. Replace YOUR_* placeholders with real values; add --scope user to install for every project.

Review the command, arguments, and environment values before installing — MCP servers run with your local permissions.

CodeRabbit
CodeRabbit
AI writes the code. CodeRabbit catches the slop.
Try For Free →
ego lite browserego lite browser
ego lite browser
Fastest browser for AI agents to run web automation tasks, always free.
Download Free life-time →
Granola, the best AI meeting recorder
Granola, the best AI meeting recorder
Notes, actions and memory. Without a meeting bot. First month 100% off.
Download for free →
CodeHealth MCP ServerCodeHealth MCP Server
CodeHealth MCP Server
Protect your code quality, stop the AI slop.
Try For Free →
belt - the only tool your agent needs
belt - the only tool your agent needs
belt cli automatically finds the best tools and skills for your agent. image, video, music, tts...
one prompt install →
AppSignal
AppSignal
Monitor with ease. Code with confidence.
Start Free Trial →
Agent, connect blockchain
Agent, connect blockchain
Connect your Claude agent to live crypto prices and trading routes via 1inch
Get the MCP →
Block distraction from your iPhone for freeBlock distraction from your iPhone for free
Block distraction from your iPhone for free
Block distracting apps from your iPhone permanently without a 3rd party app. Free and open source.
Block now (100% free) →
CodeRabbit
CodeRabbit
AI writes the code. CodeRabbit catches the slop.
Try For Free →
ego lite browserego lite browser
ego lite browser
Fastest browser for AI agents to run web automation tasks, always free.
Download Free life-time →
Granola, the best AI meeting recorder
Granola, the best AI meeting recorder
Notes, actions and memory. Without a meeting bot. First month 100% off.
Download for free →
CodeHealth MCP ServerCodeHealth MCP Server
CodeHealth MCP Server
Protect your code quality, stop the AI slop.
Try For Free →
belt - the only tool your agent needs
belt - the only tool your agent needs
belt cli automatically finds the best tools and skills for your agent. image, video, music, tts...
one prompt install →
AppSignal
AppSignal
Monitor with ease. Code with confidence.
Start Free Trial →
Agent, connect blockchain
Agent, connect blockchain
Connect your Claude agent to live crypto prices and trading routes via 1inch
Get the MCP →
Block distraction from your iPhone for freeBlock distraction from your iPhone for free
Block distraction from your iPhone for free
Block distracting apps from your iPhone permanently without a 3rd party app. Free and open source.
Block now (100% free) →

Tools

Verified live against the running server on Jun 11, 2026.

verified live10 tools
get-server-infoReturn server configuration and enabled Opik capabilities. ⚠️ DEPRECATED — migrate to `uvx opik-mcp@latest` by 2026-11-15

Return server configuration and enabled Opik capabilities. ⚠️ DEPRECATED — migrate to `uvx opik-mcp@latest` by 2026-11-15

No parameters — call it with no arguments.

get-opik-helpReturn Opik capability documentation, optionally filtered by topic. ⚠️ DEPRECATED — migrate to `uvx opik-mcp@latest` by 2026-11-151 params

Return Opik capability documentation, optionally filtered by topic. ⚠️ DEPRECATED — migrate to `uvx opik-mcp@latest` by 2026-11-15

Parameters* required
topicstring
Optional capability topic to describeone of prompts · projects · traces · metrics · general
get-opik-examplesReturn Opik usage examples for a requested task. ⚠️ DEPRECATED — migrate to `uvx opik-mcp@latest` by 2026-11-151 params

Return Opik usage examples for a requested task. ⚠️ DEPRECATED — migrate to `uvx opik-mcp@latest` by 2026-11-15

Parameters* required
taskstring
Optional task name, e.g. "create prompt", "log trace", "evaluate response"
get-opik-metrics-infoReturn Opik metric definitions and usage guidance. ⚠️ DEPRECATED — migrate to `uvx opik-mcp@latest` by 2026-11-151 params

Return Opik metric definitions and usage guidance. ⚠️ DEPRECATED — migrate to `uvx opik-mcp@latest` by 2026-11-15

Parameters* required
metricstring
Optional metric name (e.g. hallucination, answerrelevance, moderation)
get-opik-tracing-infoReturn tracing guidance for traces, spans, feedback, search, and visualization. ⚠️ DEPRECATED — migrate to `uvx opik-mcp@latest` by 2026-11-151 params

Return tracing guidance for traces, spans, feedback, search, and visualization. ⚠️ DEPRECATED — migrate to `uvx opik-mcp@latest` by 2026-11-15

Parameters* required
topicstring
Optional tracing topicone of traces · spans · feedback · search · visualization
list-projectsList projects in the active workspace to find IDs for traces and metrics operations. ⚠️ DEPRECATED — migrate to `uvx opik-mcp@latest` by 2026-11-153 params

List projects in the active workspace to find IDs for traces and metrics operations. ⚠️ DEPRECATED — migrate to `uvx opik-mcp@latest` by 2026-11-15

Parameters* required
pageinteger
1-based page number.default: 1
sizeinteger
Page size (1-100).default: 10
workspaceNamestring
Workspace override for local/stdio mode. Ignored when remote token-to-workspace mapping is enforced.
list-tracesList traces for a project for quick inspection and navigation. ⚠️ DEPRECATED — migrate to `uvx opik-mcp@latest` by 2026-11-155 params

List traces for a project for quick inspection and navigation. ⚠️ DEPRECATED — migrate to `uvx opik-mcp@latest` by 2026-11-15

Parameters* required
pageinteger
1-based page number.default: 1
projectIdstring
Optional project ID. If omitted, the first available project is used.
projectNamestring
Optional project name (alternative to projectId).
sizeinteger
Page size (1-100).default: 10
workspaceNamestring
Workspace override for local/stdio mode. Ignored when remote token-to-workspace mapping is enforced.
get-trace-by-idGet full details for a trace, including metadata and serialized input/output. ⚠️ DEPRECATED — migrate to `uvx opik-mcp@latest` by 2026-11-152 params

Get full details for a trace, including metadata and serialized input/output. ⚠️ DEPRECATED — migrate to `uvx opik-mcp@latest` by 2026-11-15

Parameters* required
traceId*string
Trace ID.
workspaceNamestring
Workspace override for local/stdio mode. Ignored when remote token-to-workspace mapping is enforced.
get-trace-statsGet aggregated trace statistics (count, tokens, cost, and duration) over time. ⚠️ DEPRECATED — migrate to `uvx opik-mcp@latest` by 2026-11-155 params

Get aggregated trace statistics (count, tokens, cost, and duration) over time. ⚠️ DEPRECATED — migrate to `uvx opik-mcp@latest` by 2026-11-15

Parameters* required
endDatevalue
projectIdstring
Optional project ID. If omitted, the first available project is used.
projectNamestring
Optional project name (alternative to projectId).
startDatestring
Date in YYYY-MM-DD format.
workspaceNamestring
Workspace override for local/stdio mode. Ignored when remote token-to-workspace mapping is enforced.
get-trace-threadsList trace threads (conversation/session groupings) or fetch one thread by ID. ⚠️ DEPRECATED — migrate to `uvx opik-mcp@latest` by 2026-11-156 params

List trace threads (conversation/session groupings) or fetch one thread by ID. ⚠️ DEPRECATED — migrate to `uvx opik-mcp@latest` by 2026-11-15

Parameters* required
pageinteger
1-based page number.default: 1
projectIdstring
Optional project ID filter.
projectNamestring
Optional project name filter.
sizeinteger
Page size (1-100).default: 10
threadIdstring
Optional thread ID. When set, returns that thread instead of paginated listing.
workspaceNamestring
Workspace override for local/stdio mode. Ignored when remote token-to-workspace mapping is enforced.

Opik MCP Server

The official Model Context Protocol (MCP) server for Opik, the open-source LLM observability and evaluation platform, built by Comet. Plug your AI host (Claude Code, Cursor, VS Code Copilot, Codex, opencode, or any MCP client) directly into your Opik workspace: read traces, log scores, and save prompt versions, all from the chat.

Built for LLM engineers who already run Opik and want to drive it from the same AI assistant they code with.

You:    "Which traces in project 'demo' failed today?"
Claude: → list(entity_type="trace", project_name="demo") → "Three traces failed…"

You:    "Score trace 7f2e… 0.9 on helpfulness with reason 'great recovery'."
Claude: → write(score.create) → done

Quick start

One command registers the server with the AI clients on your machine, installs the Opik skill pack, and verifies the connection. It needs uv and no Opik SDK:

uvx opik mcp configure

It detects Claude Code, Cursor, VS Code Copilot, Codex and opencode, and sets up the server that fits your Opik:

Your OpikServerTransportSign-in
Opik Cloud (www.comet.com)the hosted server, run by CometStreamable HTTPin the browser (OAuth); no API key
Self-hosted Comet, open-source Opikthe local server, this packagestdioenv vars; an API key only where the deployment needs one

Clients load MCP servers when a session starts, so start a new session afterwards. Without a terminal, as from a coding agent or a script, name the client: uvx opik mcp configure --ai-client claude-code (or codex, cursor, vscode, opencode). Run that way it needs an existing ~/.opik.config or OPIK_API_KEY in the environment; without either, use the commands below.

Setup guide, troubleshooting and FAQ: comet.com/docs/opik/mcp-server.


Opik Cloud: the hosted server

Comet runs the server at https://www.comet.com/opik/api/v1/mcp. Your client connects over HTTP and opens a browser sign-in the first time. There is nothing to install, no API key, and no workspace to set: the server works in the workspace you pick when you sign in. After adding it, start a new session and ask: "list my Opik projects".

Add to Cursor Install in VS Code

Claude Code

claude mcp add --scope user --transport http opik-mcp https://www.comet.com/opik/api/v1/mcp
claude mcp login opik-mcp

--scope user makes the server available in every project; without it, Claude Code registers it for the current directory only. claude mcp login opens the sign-in; /mcp → Authenticate in a session does the same. Over SSH, claude mcp login opik-mcp --no-browser prints the sign-in URL to open on your own machine; the last step needs an interactive terminal (ssh -t).

Codex

codex mcp add opik-mcp --url https://www.comet.com/opik/api/v1/mcp

add opens the sign-in; codex mcp login opik-mcp opens it again. In ~/.codex/config.toml the same server is:

[mcp_servers.opik-mcp]
url = "https://www.comet.com/opik/api/v1/mcp"

Cursor

Use the button above, or add to ~/.cursor/mcp.json:

{
  "mcpServers": {
    "opik-mcp": {
      "url": "https://www.comet.com/opik/api/v1/mcp"
    }
  }
}

VS Code Copilot

Use the button above, or add to .vscode/mcp.json in your workspace or to your user mcp.json (MCP: Open User Configuration):

{
  "servers": {
    "opik-mcp": {
      "type": "http",
      "url": "https://www.comet.com/opik/api/v1/mcp"
    }
  }
}

Other clients

Any client that takes a URL can use the hosted server:

npx add-mcp https://www.comet.com/opik/api/v1/mcp --name opik-mcp

A client that can only start local commands can reach it through npx -y mcp-remote https://www.comet.com/opik/api/v1/mcp.

Opik Cloud with an API key

Where nobody can complete the browser sign-in, as when an agent runs unattended from a script or a client has no MCP OAuth support, use an API key from comet.com/api/my/settings/ instead, with the local server pointed at Opik Cloud:

claude mcp add --scope user opik-mcp \
  --env OPIK_API_KEY="$OPIK_API_KEY" \
  --env OPIK_WORKSPACE=<workspace> \
  -- uvx opik-mcp

The other clients take the same two variables in their env block. Set OPIK_WORKSPACE to the segment after comet.com/opik/ in your Opik URL (https://www.comet.com/opik/acme-ai/projects → acme-ai). Left out, the server sends default, which Comet resolves to your account's default workspace, so reads can come from the wrong workspace without an error.


Self-hosted and open-source Opik: the local server

opik-mcp runs on your machine: the client starts it with uvx opik-mcp and talks to it over stdio. Install uv once; it fetches the package, and Python 3.13 if needed, on first use:

curl -LsSf https://astral.sh/uv/install.sh | sh   # macOS / Linux
# or: brew install uv

uvx opik-mcp reuses the copy uv has cached, so it starts quickly; uv cache clean opik-mcp makes the next start fetch the newest release. uvx opik-mcp@latest asks PyPI on every start, which adds about a second and a half. Codex waits mcp_optional_startup_grace_ms (1 s by default) for servers before it builds the first tool list, so a slower start can leave the tools out of the first turn.

Env vars point the server at your Opik:

Your OpikEnv vars
Open-source Opik on this machineOPIK_URL=http://localhost:5173/api
Open-source Opik on a serverOPIK_URL=https://<host>/api, plus OPIK_API_KEY only if the deployment adds authentication
Self-hosted Comet platform (the Opik UI is at https://<host>/opik)COMET_URL_OVERRIDE=https://<host>, OPIK_WORKSPACE, OPIK_API_KEY
  • Open source serves its API at /api and has one workspace, default, so it needs no OPIK_WORKSPACE. COMET_URL_OVERRIDE would point the server at /opik/api, which open source does not serve.
  • A self-hosted Comet has named workspaces. Set OPIK_WORKSPACE to the segment after /opik/ in your Opik URL; left out, reads come from your account's default workspace, with no error to say so.
  • A self-hosted Comet that runs the MCP OAuth server (off by default) can use the hosted flow instead, at https://<host>/opik/api/v1/mcp.
  • Substitute every value. The server refuses a placeholder workspace such as <your-workspace> or ${input:OPIK_WORKSPACE}; a placeholder URL is sent as-is and fails to connect.
  • OPIK_MCP_ANALYTICS_SOURCE="" in the env block opts a self-hosted install out of the cloud-Comet source label on telemetry events.

The examples use a local open-source Opik. For a self-hosted Comet, swap in the three variables from the table. After adding the server, start a new session and ask: "list my Opik projects".

Claude Code

claude mcp add --scope user opik-mcp --env OPIK_URL=http://localhost:5173/api -- uvx opik-mcp

On a self-hosted Comet, with the key in your shell's OPIK_API_KEY:

claude mcp add --scope user opik-mcp \
  --env COMET_URL_OVERRIDE=https://<host> \
  --env OPIK_WORKSPACE=<workspace> \
  --env OPIK_API_KEY="$OPIK_API_KEY" \
  -- uvx opik-mcp

Or edit ~/.claude.json directly:

{
  "mcpServers": {
    "opik-mcp": {
      "type": "stdio",
      "command": "uvx",
      "args": ["opik-mcp"],
      "env": {
        "OPIK_URL": "http://localhost:5173/api"
      }
    }
  }
}

claude mcp get opik-mcp shows ✔ Connected once the server starts. That does not prove the URL or key are right, because only a tool call reaches Opik. It also prints the env block, API key included.

Codex

codex mcp add opik-mcp --env OPIK_URL=http://localhost:5173/api -- uvx opik-mcp

Or edit ~/.codex/config.toml:

[mcp_servers.opik-mcp]
command = "uvx"
args = ["opik-mcp"]
env = { OPIK_URL = "http://localhost:5173/api" }
# On a self-hosted Comet, forward the key from the environment Codex starts in,
# so it stays out of this file:
# env_vars = ["OPIK_API_KEY"]
# Time allowed for the server to start (default 10 s); the first start
# downloads the package.
startup_timeout_sec = 30

Codex gives a local server only the variables in env and the names in env_vars, so a key exported in your shell does not reach it otherwise.

Cursor

Add to ~/.cursor/mcp.json (global) or .cursor/mcp.json (project), or use Cmd+Shift+J → Features → Model Context Protocol:

{
  "mcpServers": {
    "opik-mcp": {
      "type": "stdio",
      "command": "uvx",
      "args": ["opik-mcp"],
      "env": {
        "OPIK_URL": "http://localhost:5173/api"
      }
    }
  }
}

Cursor 60s timeout. Cursor enforces a hard tool-call timeout that doesn't reset on progress notifications. See Known host limits.

VS Code Copilot

Add to .vscode/mcp.json in your workspace, or to your user mcp.json (MCP: Open User Configuration):

{
  "servers": {
    "opik-mcp": {
      "type": "stdio",
      "command": "uvx",
      "args": ["opik-mcp"],
      "env": {
        "OPIK_URL": "http://localhost:5173/api"
      }
    }
  }
}

MCP Inspector (manual testing)

OPIK_URL=http://localhost:5173/api npx @modelcontextprotocol/inspector uvx opik-mcp

Install with a coding agent

For an AI agent asked to install the server. The commands are in the two sections above; these rules pick which one to run. uvx opik-mcp --help prints these rules and the commands for each case.

  1. If Opik runs on this machine — curl -s http://localhost:5173/api/is-alive/ping answers — use the local server with OPIK_URL=http://localhost:5173/api. In a sandboxed shell, such as Codex's, a failed check can mean the shell has no network access rather than that Opik is down; ask the user.
  2. If the user is on Opik Cloud (www.comet.com), use the hosted server.
  3. Otherwise ask the user for their Opik URL, and use the local server with the env vars for that deployment.
  • Don't ask the user to paste an API key into the chat. The hosted server needs none; otherwise take it from the shell environment ("$OPIK_API_KEY"), or let the user run the command.
  • Don't guess the workspace. It is the segment after comet.com/opik/ (or /opik/ on a self-hosted Comet) in the user's Opik URL.
  • An opik-mcp entry may already exist, from the old npx setup or an earlier attempt. Tell the user before replacing it. Claude Code refuses to add over it, so remove it first with claude mcp remove opik-mcp --scope user; Codex's add replaces it.
  • After registering, in Claude Code, read only the status line: claude mcp get opik-mcp | grep Status. The full output prints the env block, API key included. The hosted server shows ! Needs authentication until claude mcp login opik-mcp has run. Codex has nothing that starts the server before a session; codex mcp get opik-mcp shows what was stored, with env values masked.
  • Clients load MCP servers when a session starts. Ask the user to start a new session, then try "list my Opik projects".

Coming from npx opik-mcp?

The TypeScript server (npm opik-mcp@2) is deprecated and stops serving requests on 2026-11-15. On Opik Cloud, switch to the hosted server and drop the API key. Otherwise, in your MCP client config, replace npx -y opik-mcp with uvx opik-mcp. Some env vars were renamed and command-line flags are no longer read: see the migration guide. Support policy: DEPRECATED.md. The TypeScript source is at the git tag legacy-typescript-final.


Tools

opik-mcp exposes a small, outcome-oriented surface that covers the full lifecycle (read → annotate → curate → author → iterate).

ToolPurpose
readUniversal read by id / name / opik:// URI
listUniversal list with optional name filter + pagination
writeUniversal write — log traces/spans, score, comment, save prompts, manage datasets & experiments
schemaIntrospect write-operation schemas (used by the LLM to construct valid payloads)
read_skillRead one of the Opik agent skills bundled with this server

read

One tool for any "show me X" question. Takes an entity_type plus an id (UUID or, for nameable types, a name) or a full opik:// URI. Composite reads (trace, prompt, thread, agent_insights_issue) inline their children so a single call returns the full picture.

The record you name comes back whole. Inlined children do not: their bodies are fetched with the backend's truncate=true, so a field over ~10 KB is cut in ClickHouse and base64 images are replaced with "[image]" — one attachment echoed across 200 spans would otherwise cost more than everything else in the read. The answer says so in spanBodies / messageBodies, and any child is whole again through its own read("span", id) or read("trace", trace_id), which hit endpoints that have no truncate parameter at all.

An inlined collection is also bounded in length: 200 spans, 200 turns, 100 prompt versions. Past that, spansTruncated / messagesTruncated / versionsTruncated is true and a moreSpans / moreMessages / moreVersions line beside it carries the count and the exact list(...) call that continues from where the inlined part stopped.

Supported entities: project, trace, span, dataset, dataset_item, experiment, prompt, thread, agent_insights_issue. Name-based lookup is available for project, experiment, prompt, dataset (slower — two API calls — and may return multiple matches). thread and agent_insights_issue are project-scoped: pass project_id or project_name, or a link/URI that carries the project. dataset and dataset_item were called test_suite and test_suite_item before; the old names still resolve, but they are not advertised and new code should use the new ones.

read(entity_type="trace", id="7f2e3c8a-…")
read(entity_type="project", id="demo")  # name lookup
read(entity_type="trace", id="opik://traces/7f2e3c8a-…")
read(entity_type="agent_insights_issue", id="<issue-uuid>", project_id="<project-uuid>")
read(
    entity_type="agent_insights_issue",
    id="https://www.comet.com/opik/<ws>/projects/<pid>/diagnostics?issue=<id>",
)

A link copied from the Opik UI works as the id: a thread link or a Diagnostics page link carries the project, so no project_id is needed and the entity type is taken from the link.

A project read answers "how is my project doing" in one call. It returns {project, summary, vocabulary, contains, url}: the record, then the four figures the Logs page shows as cards (trace count, error rate, average duration, total cost) for the last 7 days against the 7 before, SDK traffic only, as on screen. since / until move that window; since="30d" is what the UI opens on. A rate or an average over a period with no traces comes back as null, because 0% errors on a week with no traffic reads as a healthy week.

vocabulary is the map you need before you can ask anything else: the project's feedback score names, its token usage keys, and the automation rules scoring its traces. These are the names that go into a filter or into series= below, and guessing them returns an empty page that reads like good news. Score names and rules are capped, always report the true total, and name the call that returns the rest; usage keys are listed in full, since nothing else enumerates them. contains names the freshest experiment, dataset, prompt version and optimization run, so "what has been happening here" does not need four more calls. A part that failed to load says so instead of looking empty, and an empty one is omitted.

An agent_insights_issue read returns {issue, example_trace_ids, details}: the Diagnostics issue record (name, description, cause, suggested fix, severity, status), the deduplicated ids of the traces that exhibit it (the same sample the Diagnostics page shows — open one with read("trace", id)), and the per-day breakdown. Trace bodies are not inlined, so the read stays one backend call. since / until narrow the per-day rows; the default is all-time. When the server knows the Opik URL and the session's workspace, the read also carries url (the issue's Diagnostics page) and trace_url_template (a deep link for any of the example traces), so the assistant can hand you something clickable; under an OAuth session whose workspace could not be resolved the links are omitted rather than guessed.

Traces themselves carry no URL — a link for one is not derivable from the fields a read or list returns, and a guessed shape 404s. The session instructions name a template for it instead, .../v1/session/redirect/projects/?trace_id={trace_id}&path=..., so the assistant fills in an id and hands you a link. It goes through opik-backend's redirect, which resolves the project and the workspace from the trace, so it works where a direct project URL cannot, an OAuth session with an unresolved workspace included. It is the same link the Python SDK prints for a trace.

list

Browse or search a collection with pagination. Project-scoped types (trace, span, thread, agent_insights_issue, dataset_item, prompt_version) need their parent: a project UUID or name, a dataset UUID, or a prompt UUID.

list(entity_type="experiment", page=1, size=25)
list(entity_type="experiment", name="rerank")  # name substring filter
list(entity_type="agent_insights_issue", project_name="demo")  # open Diagnostics issues
list(entity_type="agent_insights_issue", project_id="<uuid>", status="resolved")
list(entity_type="trace", project_name="demo")  # latest traces of one project
list(
    entity_type="trace", project_name="demo", filters="error_info is_not_empty AND duration > 5000"
)
list(
    entity_type="span",
    project_name="demo",  # spans across the whole project
    filters='type = "llm" AND usage.total_tokens > 10000',
)
list(
    entity_type="thread",
    project_name="demo",
    filters="number_of_messages > 20 AND feedback_scores.helpfulness < 0.5",
)
list(entity_type="experiment", filters='dataset_id = "<dataset-uuid>" AND tags contains "baseline"')

Filters. trace, span, thread, experiment and dataset_item take an OQL string, the same grammar as the SDK's search_traces(filter_string=…):

<field>[.<key>] <op> <value> [AND ...]
ops: = != > >= < <= contains not_contains starts_with ends_with is_empty is_not_empty in not_in

Strings go in double quotes, numbers are bare, duration is in milliseconds, dates are ISO-8601 instants with a timezone ("2026-09-08T10:00:00Z"). Scores and dictionaries take a key: feedback_scores.accuracy < 0.5, metadata.environment = "prod". AND is the only connector.

Like the UI's Logs page, trace, span and thread lists add source = "sdk" so evaluator, playground and experiment traces stay out of the way; name source yourself to see them. The first output line echoes the filter that was applied.

A bad filter fails before reaching the backend with what is needed to fix it: the position of a syntax error, the closest field name, the valid operators for the field's type, or the expected value format. Fields with a closed set of values (source, span type, thread status, visibility_mode) are checked against it too, every element of an in list included. source is the one the backend validates itself, and it answers an unknown value with a 500 rather than a 400, so source = "SDK" would otherwise be an opaque server error for a capital letter. The rest are compared as strings and answer with an empty page, which reads as "no matches" when it means "no such value". Ask schema("list.trace") (or list.span, list.thread, list.experiment) for the full field reference, accepted values included.

Finding one case in a dataset. list(entity_type="dataset_item", dataset_id=…) filters on the case itself: data.<key> for the keys the dataset was built with, full_data for a substring of the whole payload (a full scan — name a key when you can), plus id, tags, source, trace_id, span_id and the timestamps. data.<key> takes the six string operators only (=, !=, contains, not_contains, starts_with, ends_with); the backend answers a comparison with a 400, so this one is refused before the call. The endpoint has no sorting and no free-text search — sort is refused rather than dropped. read(entity_type="dataset_item", id=…) returns one case whole, which is how a value the table cut is read back.

list(entity_type="dataset_item", dataset_id="<uuid>", filters='data.question contains "install"')
list(
    entity_type="dataset_item", dataset_id="<uuid>", filters='trace_id = "<trace-uuid>"'
)  # the case made from that trace
read(entity_type="dataset_item", id="<item-uuid>")  # the case, uncut

With experiment_ids the same list is the comparison instead — the cases with each run attached — and it filters on the runs (feedback_scores.<name>, output, duration). The two are different field sets on two backend endpoints: schema("list.dataset_item_case") is the dataset's own cases, schema("list.dataset_item") the comparison.

Sort. trace, span, thread and experiment take sort="<field> [asc|desc]", desc by default and one field only: sort="duration desc", sort="total_estimated_cost", sort="feedback_scores.accuracy asc", sort="usage.total_tokens". The field is checked against the entity's sortable list before the call, because the backend silently ignores fields it cannot sort by. On very large workspaces the backend drops sorting altogether; the header says so when that happens.

dataset_item sorts only as a comparison (with experiment_ids): the items endpoint takes no sorting parameter, so a sort on a plain listing is refused rather than dropped.

Time window and search. trace, span and thread take since and until, each a relative span ("30m", "1h", "7d") or an ISO-8601 instant with a timezone, so "the last hour" needs no clock arithmetic. The window is by record creation time, which is cheap for the backend and agrees with start_time within seconds for live traffic. For an exact bound, put start_time in filters. The same three types take search, free text matched anywhere in id, name, input, output, metadata, tags and thread id. Search scans the whole project on the backend, so the first call on a large project can take tens of seconds. Those calls get a 60-second timeout. Adding since makes them fast again.

Reading the table. Durations are labelled duration_ms / ttft_ms and shown as whole milliseconds; the field stays duration in filters and sort. Timestamps are shown to the second and costs as plain decimals. Project rows carry last_updated_trace_at so you can see which project has live traffic; thread rows carry the first message. An empty page under a time window says when the project's last trace landed, and an empty page under the default source = "sdk" says how to see the other sources. A misspelled project_name comes back with the closest existing name.

list(
    entity_type="trace",
    project_name="demo",
    since="1h",
    filters="error_info is_not_empty",
    sort="duration desc",
)
list(entity_type="trace", project_name="demo", search="order-42")

Diagnostics issues. agent_insights_issue is the Diagnostics page over the MCP: the recurring failures Opik's Diagnostics job grouped for a project, ranked as the UI ranks them (most recently seen first). Columns are severity, status, total_occurrences (all-time sum), latest_count (the most recent report day, the number the issue's own description refers to) and last_seen. Open issues are listed by default; pass status="resolved" or "closed" for the rest. read and list also answer to issue, which is what the UI calls these; the long name is the one in the entity_type enum, so that one entity does not appear there twice. Counts are all-time so they match the UI; the same since / until as for traces narrow the window, truncated to UTC report days because Diagnostics aggregates per day.

An empty list says why it is empty, because "nothing is broken" and "nobody turned Diagnostics on" read the same otherwise. There are five states: Diagnostics is unavailable on this deployment, not enabled for this project, turned off, enabled but not scanned recently, or enabled and clean with the time of the last scan. The ones you can act on name the call to make, and every state links the project's Diagnostics page.

A non-empty list dates itself. The issues are whatever the last scan grouped, so the reply ends with Report covers data through <time>, and when the window you asked about runs past that, it names the uncovered tail and how to close it: a trigger when a rescan reaches back far enough, otherwise raw traces with the since it gives you. Ask for a week on a project scanned nightly and the last day is missing from the grouped answer; this is what says so.

write("agent_insights_job.enable", {"project_name": "demo"}) turns Diagnostics on. It scans daily from then on, and calling it again is safe. write("agent_insights_job.trigger", …) scans the last 24 hours now, without waiting for the nightly run. Both take the permission that reading issues takes, and both refuse where the deployment has no Diagnostics.

An issue moves through its lifecycle with write("agent_insights_issue.resolve", {"issue_id": "<uuid>", "project_name": "demo"}) — dealt with — or …close for one not worth acting on, and …reopen to put either back on the open list. All three take the same permission and answer with a link to the view the issue moved to, since a resolved issue is no longer on the default page. Whether a failure is fixed is a judgment call, so these are for when you ask: the assistant has no business tidying the list while triaging it.

Metrics over time. project_metric charts one metric for a project as a table of time buckets: trace, span and thread counts, durations, error rates, costs, token usage and feedback scores. It answers the question that follows the overview, which is when something changed.

list(entity_type="project_metric", project_name="demo", metric_type="trace_count")
list(
    entity_type="project_metric",
    project_name="demo",
    metric_type="trace_error_rate",
    since="14d",
    interval="daily",
)
list(
    entity_type="project_metric", project_name="demo", metric_type="span_count", breakdown="model"
)  # one column per model
list(
    entity_type="project_metric",
    project_name="demo",
    metric_type="span_duration",
    breakdown="model",
    series="p99",
)  # the p99 of each model

Rows are time buckets, not records, so page, size and sort are refused rather than ignored. interval is hourly, daily, weekly or total; left out, it follows the window the way the Metrics tab does — hourly up to 3 days, daily up to 30, weekly beyond — so a default chart is a few dozen rows whatever the range, and an hourly month (721 rows) is something you ask for. since / until take the same forms as everywhere else and default to the last 7 days. filters uses the fields of whichever entity the metric is about, so a span metric is filtered by span fields.

breakdown splits each bucket by tags, name, error_info, error_type, model, provider, span_type, guardrail_name or metadata.<key>. Not every metric accepts every one of those, and seven accept none at all; the tool knows which and says so before calling the backend, naming a metric that does answer the same question where one exists. Three families come back as several series at once (a duration as p50/p90/p99, a feedback score per name, token usage per key), and the backend charts one of them at a time when grouping, so series= picks it: a percentile, a score name, or a usage key. Duration defaults to p50 and token usage to total_tokens, and whichever was used is echoed on the first line.

Empty buckets are left out and counted underneath, so a quiet month is a few rows instead of a column of zeros, and a rate over a bucket with no traces is absent rather than reported as zero.

Ask schema("list.project_metric") for the metric table, the intervals and the per-metric grouping matrix.

A project's names. score_name lists the feedback score names recorded in a project and online_rule the automation rule evaluators configured on it, which is where most of those names come from. Both are the same lists read("project", …) carries, in full and paginated, for when the capped version in the overview is not enough.

list(entity_type="score_name", project_name="demo")
list(entity_type="online_rule", project_name="demo")

write

Universal write dispatcher. Pass operation + data and the dispatcher validates the payload, applies the right REST verb, and returns the backend response.

Operations:

OperationWhat it does
trace.createLog a single trace (or a batch). Parent for spans / scores / comments.
trace.updateFinalize or amend an existing trace.
span.createLog a span on an existing trace (or a batch).
score.createAttach a numeric feedback score to a trace, span, or thread.
comment.createAttach a free-text comment to a trace, span, or thread.
prompt_version.saveSave a new prompt version (creates the prompt by name if missing).
dataset.createCreate a dataset — type: "test_suite" makes it an evaluation test suite.
dataset_item.upsertUpsert items into a dataset (always the envelope shape).
experiment.createCreate an experiment scoped to a dataset.
experiment_item.createAttach trace + dataset_item rows to an experiment.
thread.closeClose a thread (mark it inactive). Pass thread_id and the project.
thread.openReopen a closed thread. Pass thread_id and the project.
agent_insights_job.enableTurn Diagnostics on for a project (daily scans, safe to repeat).
agent_insights_job.triggerRun a Diagnostics scan now, over the last 24 hours.
agent_insights_issue.resolveMark a Diagnostics issue dealt with (ask the user first).
agent_insights_issue.closeMark a Diagnostics issue not worth acting on (ask the user first).
agent_insights_issue.reopenPut a resolved or closed Diagnostics issue back on the open list.
write(
    operation="score.create",
    data={
        "target": "trace",
        "target_id": "7f2e3c8a-…",
        "name": "helpfulness",
        "value": 0.9,
        "reason": "great recovery",
    },
)

schema

Inspect the exact JSON shape and required fields of any write operation before you call it — useful when you're not sure what data should look like. Returns the schema, OAuth scope, and one validated example. Pure lookup, no backend call.

schema(operation="score.create")
schema(operation="prompt_version.save")

The same tool answers list.trace, list.span, list.thread and list.experiment with the list tool's reference for that entity: every filterable field with its type and valid operators, the sortable fields, whether a time window and free-text search apply, and two example filters.

schema(operation="list.trace")

Configuration

These configure the local server; every setting is an environment variable. The hosted server on Opik Cloud takes none of them.

Identity / endpoint

VariableDefaultNotes
OPIK_API_KEY—API key, for a self-hosted Comet, or for Opik Cloud without the hosted server. Open-source Opik needs none unless the deployment adds authentication.
OPIK_WORKSPACEunsetWorkspace name. On cloud with an API key, unset sends default, which resolves to your account's default workspace — set it explicitly if you work in a different one, or reads come from the wrong workspace silently. Leave unset over OAuth (the token carries it) and on local/OSS (default is the only workspace there).
COMET_WORKSPACE—Deprecated alias for OPIK_WORKSPACE (backward compat). OPIK_WORKSPACE wins if both are set.
COMET_WORKSPACE_IDunsetOptional workspace UUID. Stamped into analytics events when set, and takes precedence over the resolved one. Rarely needed — OAuth installs get the UUID from the token automatically.
COMET_URL_OVERRIDEhttps://www.comet.comSet to your self-hosted Comet host, or https://dev.comet.com for staging.
OPIK_URLderived from COMET_URL_OVERRIDE + /opik/apiSet it for open-source Opik, which serves its API at /api (http://localhost:5173/api locally). On a Comet platform, override only if Opik lives on a different host/path than the Comet UI.
OPIK_DEFAULT_PROJECT_NAMEunsetWhen set, the per-session instructions blob tells the LLM to pass this as project_name on every tool call unless the user names a different project.

Server / transport

VariableDefaultNotes
OPIK_MCP_TRANSPORTstdiostdio for host-launched, streamable-http to listen on a port.
OPIK_MCP_HOST127.0.0.1uvicorn bind host (streamable-http only).
OPIK_MCP_PORT8080uvicorn bind port (streamable-http only).
OPIK_MCP_RELOADfalsetrue to enable uvicorn --reload (dev only).
OPIK_MCP_AS_URLunsetOAuth Authorization Server URL, advertised in /.well-known/oauth-protected-resource (RFC 9728) and used as the proxy target for AS-discovery probes. Required for MCP hosts to bootstrap the OAuth dance over HTTP.
OPIK_MCP_RESOURCE_URIunsetCanonical public URI of this server, advertised as resource in the protected-resource metadata and used to derive the WWW-Authenticate hint.
OPIK_MCP_OAUTH_VALIDATION_CACHE_TTL_S30How long a "valid" answer from opik-backend's token introspection is trusted before the next request on the same OAuth token asks again. Bounds the backend load added by per-request validation and the window in which an expired token is still forwarded (that window also ends on the first 401 the backend returns). Capped by the token's own expires_at when the backend reports one.
OPIK_MCP_LOG_LEVELINFOstderr logger threshold.
Choosing a transport

Two bearer shapes, two contracts on HTTP transport. An opik_mcp_at_… OAuth access token is validated on every request against opik-backend's token introspection endpoint (cached, see OPIK_MCP_OAUTH_VALIDATION_CACHE_TTL_S); an expired or revoked token gets an HTTP 401 with WWW-Authenticate: Bearer error="invalid_token", which is what MCP hosts key their silent refresh_token grant on. An Opik API key is not validated locally: it is forwarded verbatim to opik-backend, which is its single point of enforcement. Pick the transport by deployment shape:

ScenarioTransport
Opik CloudNothing to run: the hosted server is this server over HTTP, run by Comet
MCP client and Opik on the same machine (local OSS install)stdio (recommended — simplest, no port, no OAuth setup)
Local MCP client → self-hosted Opikstdio with the env vars for the deployment, or HTTP with OAuth (OPIK_MCP_AS_URL pointing at the backend)
opik-mcp served behind the same edge as opik-backendHTTP — bearers are validated by the backend per request

Note for local OSS installs: the OSS backend does not authenticate requests, so an HTTP opik-mcp in front of it is as open as the OSS REST API itself. Keep the default 127.0.0.1 bind (and prefer stdio) on shared networks.

Telemetry

Anonymous usage events (event type + timing only — no query content). A SHA-256 digest of your API key is included so support can find your account; the raw key never leaves the process. Opt out: OPIK_MCP_ANALYTICS_ENABLED=false.

VariableDefaultNotes
OPIK_MCP_ANALYTICS_ENABLEDtrueSet to false to disable all telemetry.
OPIK_MCP_ANALYTICS_URLhttps://stats.comet.com/notify/event/Override for staging.
OPIK_MCP_ANALYTICS_ENVIRONMENTprodTag on every event (prod / staging / dev).
OPIK_MCP_ANALYTICS_SOURCEcomet.comReceiver uses this to mark on_prem=False. On-prem installs should override to "" or their own domain.
OPIK_MCP_ANALYTICS_CONNECT_TIMEOUT_S5.0HTTP connect timeout.
OPIK_MCP_ANALYTICS_TOTAL_TIMEOUT_S10.0HTTP total request timeout.

Known host limits

Hosts differ in how long they let a single tool call run:

  • Claude Code — no documented tool-call timeout. Recommended.
  • Cursor — hard 60s timeout that does not reset on progress (upstream bug).
  • MCP Inspector — MAX_TOTAL_TIMEOUT bounds total duration (default 60s). Raise it in the Inspector UI for long operations.

If a call gets stuck, set OPIK_MCP_LOG_LEVEL=DEBUG for the full request log.


Troubleshooting

OPIK_API_KEY isn't picked up — the var isn't reaching the server process. In Claude Code / Cursor / VS Code, env vars only apply when inside the env block of the MCP server config, not your shell; Codex also forwards the names listed in env_vars. Start a new session after editing, since clients read the config when a session starts.

Requests go to /opik/api on an open-source Opik — COMET_URL_OVERRIDE is for a self-hosted Comet platform. Open source serves its API at /api: set OPIK_URL=http://localhost:5173/api (or https://<host>/api) instead.

Cursor call times out at 60s — Cursor's known bug, not opik-mcp. Either narrow the call (smaller size, a tighter window), or run the same operation on Claude Code which has no hard cap.

Server not showing, sign-in not opening, wrong workspace, uvx not found. These are covered in the troubleshooting section of the docs. opik mcp status (from the same uvx opik CLI) lists every client that has the server configured and whether its config has drifted.


Development

git clone git@github.com:comet-ml/opik-mcp.git
cd opik-mcp
make install        # uv sync --locked --extra dev
make check          # lint + typecheck + test
make run-dev        # uvicorn with --reload + DEBUG logs
make inspect        # MCP Inspector against the running server

Common targets:

TargetWhat it does
make installuv sync --locked --extra dev
make runRun the MCP server (stdio by default).
make run-devRun with DEBUG logging + uvicorn --reload.
make devRun via mcp dev (Inspector dev-mode wrapper).
make inspectLaunch MCP Inspector against a running server.
make testuv run pytest -q.
make lintruff check + format check.
make formatruff format + ruff check --fix.
make typecheckmypy.
make checklint + typecheck + test.

Repo layout:

opik-mcp/
├── src/opik_mcp/        ← server, tools, analytics
├── tests/               ← pytest suites
├── scripts/             ← live-BE smoke + MCP-session smoke
├── legacy/typescript/   ← migration guide for the deprecated v2 TS server (source: tag `legacy-typescript-final`)
├── pyproject.toml
└── Makefile

Get help

  • Open an issue for bugs and feature requests
  • Opik docs for SDK / backend documentation
  • Comet community Slack for questions

License

Apache-2.0.

Featured
CodeRabbit
CodeRabbit
AI writes the code. CodeRabbit catches the slop.
Try For Free →
ego lite browserego lite browser
ego lite browser
Fastest browser for AI agents to run web automation tasks, always free.
Download Free life-time →
Granola, the best AI meeting recorder
Granola, the best AI meeting recorder
Notes, actions and memory. Without a meeting bot. First month 100% off.
Download for free →
CodeHealth MCP ServerCodeHealth MCP Server
CodeHealth MCP Server
Protect your code quality, stop the AI slop.
Try For Free →
belt - the only tool your agent needs
belt - the only tool your agent needs
belt cli automatically finds the best tools and skills for your agent. image, video, music, tts...
one prompt install →
AppSignal
AppSignal
Monitor with ease. Code with confidence.
Start Free Trial →
Agent, connect blockchain
Agent, connect blockchain
Connect your Claude agent to live crypto prices and trading routes via 1inch
Get the MCP →
Block distraction from your iPhone for freeBlock distraction from your iPhone for free
Block distraction from your iPhone for free
Block distracting apps from your iPhone permanently without a 3rd party app. Free and open source.
Block now (100% free) →

Configuration

OPIK_API_KEY*secret

API key from your Opik workspace for authenticating SDK calls.

OPIK_API_BASE_URL

Override the API base URL when using a self-hosted Opik deployment.

OPIK_WORKSPACE_NAME

Default workspace to scope prompt and trace operations.

Registryactive
Packageopik-mcp
TransportSTDIO
AuthRequired
Resources2
Prompts2
Tools verifiedJun 11, 2026
UpdatedMay 29, 2026
View on GitHub