
Consult7 is an MCP server that enables AI agents to offload analysis of large file collections to high-capacity language models via OpenRouter, supporting context windows up to 2M tokens. It collects files from specified paths, assembles them into a single context, sends them to a selected model (such as Google's Gemini 3.1 family or other frontier models) with a user-provided query, and returns the results back to the agent. This solves the problem of agents running out of context when analyzing large codebases, document repositories, or mixed content that exceed their current limits, while also enabling access to specialized model capabilities and comparative analysis across different models.
Consult7 is a Model Context Protocol (MCP) server that enables AI agents to consult large context window models via OpenRouter for analyzing extensive file collections - entire codebases, document repositories, or mixed content that exceed the current agent's context limits.
Consult7 enables any MCP-compatible agent to offload file analysis to large context models (up to 2M tokens). Useful when:
"For Claude Code users, Consult7 is a game changer."
Consult7 collects files from the specific paths you provide (with optional wildcards in filenames), assembles them into a single context, and sends them to a large context window model along with your query. The result is directly fed back to the agent you are working with.
["/Users/john/project/src/*.py", "/Users/john/project/lib/*.py"]"google/gemini-3-flash-preview""fast"["/Users/john/webapp/src/*.py", "/Users/john/webapp/auth/*.py", "/Users/john/webapp/api/*.js"]"anthropic/claude-opus-4.8""think"["/Users/john/project/src/*.py", "/Users/john/project/tests/*.py"]"google/gemini-3.1-pro-preview""think""/Users/john/reports/code_review.md""Result has been saved to /Users/john/reports/code_review.md" plus a one-line metadata footer, instead of flooding the agent's contextConsult7 supports Google's Gemini 3.1 family:
google/gemini-3.1-pro-preview) - Flagship reasoning model, 1M contextgoogle/gemini-3-flash-preview) - Ultra-fast model, 1M contextgoogle/gemini-3.1-flash-lite-preview) - Ultra-fast lite model, 1M contextQuick mnemonics for power users:
gemt = Gemini 3.1 Pro + think (flagship reasoning)gemf = Gemini 3 Flash + fast (ultra fast)gptt = GPT-6 Astra + think (latest GPT, effort xhigh)grot = Grok 4.7 + think (effort xhigh)oput = Claude Opus 4.8 + think (adaptive thinking)fabt = Claude Fable 5.1 + think (deepest reasoning, effort xhigh; premium)ULTRA = Run GPTT, GROT, and FABT in parallel (3 frontier models)FUSE = Fusion: a frontier panel deliberates and a judge synthesizes, in one callThese mnemonics make it easy to reference model+mode combinations in your queries.
Note on Fable 5.1.
anthropic/claude-fable-5.1is Anthropic's most capable model but priced at a premium (~2× Opus 4.8). It does not replace Opus 4.8 as the everyday Claude choice for single calls. Since v3.11.0 it holds the Anthropic seat in theULTRApanel (which is meant for hard questions anyway). Unlike Opus 4.8 (adaptive thinking only), OpenRouter honors Fable's effort scale, somid/thinkmap toeffort=high/effort=xhigh.
Consult7 supports OpenRouter's Fusion (openrouter/fusion) — a single call where a panel of frontier models (Opus, GPT, Gemini Pro) answers your query in parallel and a judge model synthesizes their responses into one answer. Reach for it on hard questions where multiple perspectives help and the cost of being wrong outweighs a few extra completions.
fast / mid / think map the panel's web-search/fetch budget to max_tool_calls of 2 / 8 / 16.FUSE = openrouter/fusion.Trivial prompts answer directly (no panel); the panel fires only when the question warrants deliberation. Fusion is billed per panel run, so it costs more than a single-model call.
Simply run:
claude mcp add -s user consult7 uvx -- consult7 your-openrouter-api-key
Add to your Claude Desktop configuration file:
{
"mcpServers": {
"consult7": {
"type": "stdio",
"command": "uvx",
"args": ["consult7", "your-openrouter-api-key"]
}
}
}
Replace your-openrouter-api-key with your actual OpenRouter API key.
No installation required - uvx automatically downloads and runs consult7 in an isolated environment.
uvx consult7 <api-key> [--test]
<api-key>: Required. Your OpenRouter API key--test: Optional. Test the API connectionThe model and mode are specified when calling the tool, not at startup.
Consult7 supports all 500+ models available on OpenRouter. Below are the flagship models with optimized dynamic file size limits:
| Model | Context | Use Case |
|---|---|---|
openai/gpt-6-astra | 1M | Latest top-tier GPT, effort-based reasoning; premium price |
google/gemini-3.1-pro-preview | 1M | Flagship reasoning model |
google/gemini-3-flash-preview | 1M | Gemini 3 Flash, ultra fast |
google/gemini-3.1-flash-lite-preview | 1M | Ultra-fast lite model |
anthropic/claude-fable-5.1 | 1M | Most capable; premium price — reserved for hard problems |
anthropic/claude-opus-4.8 | 1M | Best quality, adaptive thinking |
anthropic/claude-sonnet-4.6 | 1M | Excellent reasoning, fast |
anthropic/claude-haiku-4.5 | 200k | Budget, very fast |
x-ai/grok-4.7 | 500k | Frontier Grok, effort-based reasoning |
x-ai/grok-4.20 | 2M | Automatic reasoning, huge context |
x-ai/grok-4.1-fast | 2M | Largest context window |
openrouter/fusion | 128k | Multi-model panel + judge (see Featured: Fusion) |
Superseded IDs still work with their tuned settings: openai/gpt-5.6-sol, x-ai/grok-4.6, anthropic/claude-fable-5.
Quick mnemonics:
gptt = openai/gpt-6-astra + think (latest GPT, deep reasoning [effort xhigh]; premium)gemt = google/gemini-3.1-pro-preview + think (Gemini 3.1 Pro, flagship reasoning)grot = x-ai/grok-4.7 + think (Grok 4.7, deep reasoning [effort xhigh]; 500K context — use x-ai/grok-4.20 for bigger bundles)oput = anthropic/claude-opus-4.8 + think (Claude Opus, adaptive thinking)opuf = anthropic/claude-opus-4.8 + fast (Claude Opus, no reasoning)fabt = anthropic/claude-fable-5.1 + think (Claude Fable, deepest reasoning [effort xhigh]; premium, hard problems only)fabm = anthropic/claude-fable-5.1 + mid (Claude Fable, high-effort reasoning; premium)gemf = google/gemini-3-flash-preview + fast (Gemini 3 Flash, ultra fast)ULTRA = call GPTT, GROT, and FABT IN PARALLEL (3 frontier models for maximum insight)FUSE = openrouter/fusion (one call: a frontier panel deliberates, a judge synthesizes; mode sets web-research depth)You can use any OpenRouter model ID (e.g., deepseek/deepseek-r1-0528). See the full model list. File size limits are automatically calculated based on each model's context window.
fast: No reasoning requested - quick answers, simple tasks. GPT-6 Astra, Grok 4.7 and Fable reason by design, so on them fast means their own default level (billed), shown in the footer as reasoning: model defaultmid: Moderate reasoning - code reviews, bug analysisthink: Maximum reasoning - security audits, complex refactoring/Users/john/project/src/*.py/Users/john/project/*.py (not in directory paths)*.py not *["/path/src/*.py", "/path/README.md", "/path/tests/*_test.py"]Common patterns:
/path/to/dir/*.py/path/to/tests/*_test.py or /path/to/tests/test_*.py["/path/*.js", "/path/*.ts"]Automatically ignored: __pycache__, .env, secrets.py, .DS_Store, .git, node_modules. Wildcards skip them; naming one explicitly is an error.
Fail fast: a relative or missing path, a directory, a wildcard that matches nothing, an explicitly named ignored file, a binary file (NUL bytes in the first 8 KB), or files over the model's size budget fail the whole call before anything is sent to the model (no cost). The error lists every problem.
Size budget: one total for all files together, no per-file limit: (model context − output reserve − reasoning reserve) × ~4 bytes per token, e.g. ~4 MB for 1M-context models, ~2 MB for Grok 4.7 (500K), ~8 MB for Grok 4.20 (2M). A second check on the estimated token count keeps a 10% safety margin. A size error states the budget and the total requested.
The consultation tool accepts the following parameters:
fast, mid, or think_updated suffix (e.g., report.md → report_updated.md, then report_updated_1.md, ...)"Result has been saved to /path/to/file" plus the metadata footerfalse)
true, routes only to endpoints with ZDR policy (prompts not retained by provider)Claude Code will automatically use the tool with proper parameters:
{
"files": ["/Users/john/project/src/*.py"],
"query": "Explain the main architecture",
"model": "google/gemini-3-flash-preview",
"mode": "fast"
}
from consult7.consultation import consultation_impl
result = await consultation_impl(
files=["/path/to/file.py"],
query="Explain this code",
model="google/gemini-3-flash-preview",
mode="fast", # fast, mid, or think
provider="openrouter",
api_key="sk-or-v1-..."
)
# Test OpenRouter connection
uvx consult7 sk-or-v1-your-api-key --test
To remove consult7 from Claude Code:
claude mcp remove consult7 -s user
model:nitro use the base model's context size.time, the mode always shown ([fast] too), zdr shown when on. Both token estimates now use the same prompt text, and a rejected fast call no longer claims reasoning disabled..env, secrets.py, ...), or files over the model's per-file or total size limit now fail the call with a clear error before anything is sent. Before, these were reported only inside the prompt (or dropped), and the paid call went through on partial input.output_file is checked before the call. A relative or unwritable path fails at no cost. If saving still fails after the call, the response is returned instead of being lost.isError=true in the MCP result, so clients and wrappers can detect failures without parsing text.user_id.1 file (not 1 files). On fast, models that always reason (GPT-6 Astra, Grok 4.7, Fable) show reasoning: model default.gemt and all Gemini models stay available); Opus 4.8 (oput/opuf) stays available but its ULTRA seat goes to Fable.gptt → GPT-6 Astra (openai/gpt-6-astra, 1M context, $10/$50 per M) — also the model used by consult7 <key> --test. Reasoning is mandatory on Astra, so mid/think now map to effort=high/effort=xhigh (GPT-5.6 Sol used medium/high). ZDR supported.grot → Grok 4.7 (x-ai/grok-4.7, 500K context; grok 4.6 had been the default since v3.10.0). Same effort mapping (high/xhigh). ZDR supported.fabt/fabm → Claude Fable 5.1 (anthropic/claude-fable-5.1, 1M context). Same effort mapping. ZDR not supported.openai/gpt-5.6-sol, x-ai/grok-4.6, anthropic/claude-fable-5) keep working with their previous settings.openai/gpt-5.6-sol) — the latest top-tier GPT, ~1M context / 128K output, effort-based reasoning (mid → effort=medium, think → effort=high). Replaces GPT-5.5 as the gptt default; GPT-5.5 stays available as a legacy model. ZDR is not supported on GPT-5.6 Sol (GPT-5.5 still is).x-ai/grok-4.5 is region-restricted on OpenRouter (returns a 403 "not available in your region") and could not be verified against the real API, so it was not integrated. Grok 4.20 remains the grot default.anthropic/claude-fable-5) — Anthropic's most capable model, 1M context. Premium price (~2× Opus 4.8), so it's reserved for specifically hard problems and is not part of the ULTRA panel; it does not replace Opus 4.8 as the default Claude model. New mnemonics fabt (think) / fabm (mid). Unlike Opus 4.8 (adaptive thinking only), OpenRouter honors Fable's effort scale, so mid/think map to effort=high/effort=xhigh (max intentionally not exposed — it tends to overthink at ~2× token cost). ZDR not supported (Fable requires 30-day retention).openrouter/fusion) — a multi-model panel plus a judge in one call; mode maps to web-research depth (fast/mid/think → max_tool_calls 2/8/16). New FUSE mnemonic.oput/opuf now point to 4.8, and 4.7 is kept as a legacy ID.cost: $0.0923.mid vs think for adaptive models (Opus, Grok)output_file return now includes the metadata footer so callers can verify what ranreasoning.enabled=truereasoning.enabled=truegptt → GPT-5.5, oput/opuf → Claude Opus 4.7, grot → Grok 4.20gemt → Gemini 3.1 Pro, oput/opuf → Claude Opus 4.6google/gemini-3-flash-preview (Gemini 3 Flash, ultra fast)gemf mnemonic to use Gemini 3 Flashzdr parameter for Zero Data Retention routinggoogle/gemini-3-pro-preview (1M context, flagship reasoning model)gemt (Gemini 3 Pro), grot (Grok 4), ULTRA (parallel execution)|thinking suffix - use mode parameter instead (now required)mode parameter API: fast, mid, thinkconsult7 <provider> <key> to consult7 <key>output_file parameter to save responses to filesMIT