Gemini Mcp

2038 toolsauthSTDIOregistry active

Summary

Connects Claude to Google's Gemini 3 models through MCP, giving you access to Gemini's text generation, image creation with Nano Banana Pro, video generation via Veo 2.0, and code execution capabilities. You can run collaborative brainstorming sessions between Claude and Gemini, analyze YouTube videos and documents, perform real-time web searches with citations, and execute Python code with data visualization. The server handles multi-turn image editing sessions and supports structured JSON output with schema validation. Reach for this when you need Gemini's specific strengths like video generation or want to combine both AI systems for complex research and analysis tasks.

CodeRabbit

AI writes the code. CodeRabbit catches the slop.

Try For Free →

Make your agent a DeFi expert

Agent, run crypto. Access onchain data & trade routes via 1inch.

Install now →

Make money from your Skills

On Capafy, your Skill runs online 24/7 as an agent product, and you get paid every time someone uses it.

Start earning →

AppSignal

Monitor with ease. Code with confidence.

Start Free Trial →

Vibe Prospecting MCP

Connect Claude to +800M contacts, +150M companies. Find & Enrich leads in chat.

Try For Free →

Context.dev

Integrate web data into your AI product. One API to scrape website & brand data.

Get API Key Now →

CodeRabbit

AI writes the code. CodeRabbit catches the slop.

Try For Free →

Make your agent a DeFi expert

Agent, run crypto. Access onchain data & trade routes via 1inch.

Install now →

Make money from your Skills

On Capafy, your Skill runs online 24/7 as an agent product, and you get paid every time someone uses it.

Start earning →

AppSignal

Monitor with ease. Code with confidence.

Start Free Trial →

Vibe Prospecting MCP

Connect Claude to +800M contacts, +150M companies. Find & Enrich leads in chat.

Try For Free →

Context.dev

Integrate web data into your AI product. One API to scrape website & brand data.

Get API Key Now →

Tools

Public tool metadata for what this MCP can expose to an agent.

8 tools

GEMINI_COUNT_TOKENSCounts the number of tokens in text using Gemini tokenization. Useful for estimating costs, checking input limits, and optimizing prompts before making API calls.2 params

Counts the number of tokens in text using Gemini tokenization. Useful for estimating costs, checking input limits, and optimizing prompts before making API calls.

Parameters* required

textstring

Text to count tokens for

modelstring

Model to use for token counting. Examples: 'gemini-1.5-flash', 'gemini-1.5-pro'default: gemini-1.5-flash

GEMINI_EMBED_CONTENTGenerates text embeddings using Gemini embedding models. Converts text into numerical vectors for semantic search, similarity comparison, clustering, and classification tasks.4 params

Generates text embeddings using Gemini embedding models. Converts text into numerical vectors for semantic search, similarity comparison, clustering, and classification tasks.

Parameters* required

textstring

Text to generate embeddings for

modelstring

Embedding model to use. Examples: 'text-embedding-004', 'embedding-001'default: text-embedding-004

titlestring

Optional title for the content (for document embeddings)

task_typestring

Task type: 'RETRIEVAL_QUERY', 'RETRIEVAL_DOCUMENT', 'SEMANTIC_SIMILARITY', 'CLASSIFICATION', 'CLUSTERING'

GEMINI_GENERATE_CONTENTGenerates text content from prompts using Gemini models. Supports various models like Gemini Flash and Pro with configurable temperature, token limits, and safety settings for diverse text generation tasks.9 params

Generates text content from prompts using Gemini models. Supports various models like Gemini Flash and Pro with configurable temperature, token limits, and safety settings for diverse text generation tasks.

Parameters* required

modelstring

Model to use. Examples: 'gemini-1.5-flash', 'gemini-1.5-pro', 'gemini-2.0-flash-exp'default: gemini-1.5-flash

top_kinteger

Top-k sampling parameter

top_pnumber

Nucleus sampling parameter (0.0 to 1.0)

promptstring

Text prompt for content generation

temperaturenumber

Controls randomness (0.0 to 2.0)

stop_sequencesarray

Sequences where generation should stop

safety_settingsarray

Safety filter settings

max_output_tokensinteger

Maximum number of tokens to generate

system_instructionstring

System instruction to guide the model's behavior

GEMINI_GENERATE_IMAGEGenerates images from text prompts using Gemini 2.5 Flash Image Preview model (Nano Banana). Supports creative image generation with customizable parameters like aspect ratio, safety settings, and optional local file saving. Generated images are automatically uploaded to S3 an...9 params

Generates images from text prompts using Gemini 2.5 Flash Image Preview model (Nano Banana). Supports creative image generation with customizable parameters like aspect ratio, safety settings, and optional local file saving. Generated images are automatically uploaded to S3 an...

Parameters* required

modelstring

Model to use. Use 'gemini-2.5-flash-image-preview' for image generationdefault: gemini-2.5-flash-image-preview

top_kinteger

Top-k sampling parameter

top_pnumber

Nucleus sampling parameter (0.0 to 1.0)

promptstring

Text prompt for image generation

save_pathstring

Optional local path to save the generated image

temperaturenumber

Controls randomness (0.0 to 2.0)

safety_settingsarray

Safety filter settings

max_output_tokensinteger

Maximum number of tokens to generate (max 32,768)

system_instructionstring

System instruction to guide image generation behavior

GEMINI_GENERATE_VIDEOSGenerates videos from text prompts using Google's Veo models. Creates high-quality video content. Returns operation ID for tracking progress. After this, call GEMINI_WAIT_FOR_VIDEO to download the video using the operation ID.4 params

Generates videos from text prompts using Google's Veo models. Creates high-quality video content. Returns operation ID for tracking progress. After this, call GEMINI_WAIT_FOR_VIDEO to download the video using the operation ID.

Parameters* required

modelstring

Model to use. Examples: 'veo-3.0-generate-preview', 'veo-3.0-fast-generate-preview', 'veo-2.0-generate-001'default: veo-3.0-generate-preview

extrasobject

Additional parameters passed through to API

promptstring

Text prompt for Veo video generation

person_generationstring

Controls person generation in videos. Values: 'allow_adult' or 'dont_allow'. IMPORTANT: Veo 3 models in EU/UK/CH/MENA regions ONLY support 'allow_adult'. Veo 2 models support both values in all regions.

GEMINI_GET_VIDEOS_OPERATIONChecks the status of a Veo video generation operation. Use the operation name from GenerateVideos to track progress and get the download URL when complete.1 params

Checks the status of a Veo video generation operation. Use the operation name from GenerateVideos to track progress and get the download URL when complete.

Parameters* required

operation_namestring

Operation resource name returned by predictLongRunning

GEMINI_LIST_MODELSLists available Gemini and Veo models with their capabilities and limits. Useful for discovering supported models and their features before making generation requests.1 params

Lists available Gemini and Veo models with their capabilities and limits. Useful for discovering supported models and their features before making generation requests.

Parameters* required

filter_prefixstring

Filter models by name prefix (client-side). Leave empty to get all models.default:

GEMINI_WAIT_FOR_VIDEOPolls a Veo video generation operation until completion, then downloads and returns the video as a FileDownloadable with public URL.1 params

Polls a Veo video generation operation until completion, then downloads and returns the video as a FileDownloadable with public URL.

Parameters* required

operation_namestring

The operation name from video generation (e.g., 'models/...')

MCP Server Gemini

A Model Context Protocol (MCP) server for integrating Google's Gemini 3 models with Claude Code, enabling powerful collaboration between both AI systems. Now with a beautiful CLI!

MCP Registry Support: Now discoverable in the official MCP ecosystem!

Features

Feature	Description
Deep Research Agent	Autonomous multi-step research with web search and citations
Token Counting	Count tokens and estimate costs before API calls
Text-to-Speech	30 unique voices, single speaker or two-speaker dialogues
URL Analysis	Analyze, compare, and extract data from web pages
Context Caching	Cache large documents for efficient repeated queries
YouTube Analysis	Analyze videos by URL with timestamp clipping
Document Analysis	PDFs, DOCX, spreadsheets with table extraction
4K Image Generation	Generate images up to 4K with 10 aspect ratios
Multi-Turn Image Editing	Iteratively refine images through conversation
Video Generation	Create videos with Veo 2.0 (async with polling)
Code Execution	Gemini writes and runs Python code (pandas, numpy, matplotlib)
Google Search	Real-time web information with inline citations
Structured Output	JSON responses with schema validation
Data Extraction	Extract entities, facts, sentiment from text
Thinking Levels	Control reasoning depth (minimal/low/medium/high)
Direct Query	Send prompts to Gemini 3 Pro/Flash models
Brainstorming	Claude + Gemini collaborative problem-solving
Code Analysis	Analyze code for quality, security, performance
Summarization	Summarize content at different detail levels

Quick Installation

MCP Server for Claude Code

# Using npm (Recommended)
claude mcp add gemini -s user -- env GEMINI_API_KEY=YOUR_KEY npx -y @rlabs-inc/gemini-mcp

# Using bun
claude mcp add gemini -s user -- env GEMINI_API_KEY=YOUR_KEY bunx @rlabs-inc/gemini-mcp

CLI (Global Install)

# Install globally
npm install -g @rlabs-inc/gemini-mcp

# Set your API key once (stored securely)
gcli config set api-key YOUR_KEY

# Now use any command!
gcli search "latest news"
glci image "sunset over mountains" --ratio 16:9

Get your API key: Visit Google AI Studio - it's free and takes seconds!

Installation Options

# With verbose logging
claude mcp add gemini -s user -- env GEMINI_API_KEY=YOUR_KEY VERBOSE=true bunx -y @rlabs-inc/gemini-mcp

# With custom output directory for generated images/videos
claude mcp add gemini -s user -- env GEMINI_API_KEY=YOUR_KEY GEMINI_OUTPUT_DIR=/path/to/output bunx -y @rlabs-inc/gemini-mcp

Available Tools

gemini-query

Direct queries to Gemini with thinking level control:

prompt: "Explain quantum entanglement"
model: "pro" or "flash"
thinkingLevel: "low" | "medium" | "high" (optional)

low: Fast responses, minimal reasoning
medium: Balanced (Flash only)
high: Deep reasoning for complex tasks (default)

gemini-generate-image

Generate images with Nano Banana Pro (Claude can SEE them!):

prompt: "a futuristic city at sunset"
style: "cyberpunk" (optional)
aspectRatio: "16:9" (1:1, 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9)
imageSize: "2K" (1K, 2K, 4K)
useGoogleSearch: false (ground in real-world info)
thinkingLevel: "high" (optional - minimal, low, medium, high)
personGeneration: "ALLOW_ALL" (optional - ALLOW_ALL, ALLOW_ADULT, ALLOW_NONE)
seed: 42 (optional - for reproducible results)

gemini-start-image-edit

Start a multi-turn image editing session:

prompt: "a cozy cabin in the mountains"
aspectRatio: "16:9"
imageSize: "2K"
useGoogleSearch: false
thinkingLevel: "high" (optional - minimal, low, medium, high)
personGeneration: "ALLOW_ALL" (optional - ALLOW_ALL, ALLOW_ADULT, ALLOW_NONE)
seed: 42 (optional - for reproducible results)

Returns a session ID for iterative editing.

gemini-continue-image-edit

Continue refining an image:

sessionId: "edit-123456789"
prompt: "add snow on the roof and make it nighttime"

gemini-end-image-edit

Close an editing session:

sessionId: "edit-123456789"

gemini-list-image-sessions

List all active editing sessions.

gemini-generate-video

Generate videos using Veo:

prompt: "a cat playing piano"
aspectRatio: "16:9" (optional)
negativePrompt: "blurry, text" (optional)

Video generation is async (takes 1-5 minutes). Use gemini-check-video to poll.

gemini-check-video

Check video generation status and download when complete:

operationId: "operations/xxx-xxx-xxx"

gemini-analyze-code

Analyze code for issues:

code: "function foo() { ... }"
language: "typescript" (optional)
focus: "quality" | "security" | "performance" | "bugs" | "general"

gemini-analyze-text

Analyze text content:

text: "Your text here..."
type: "sentiment" | "summary" | "entities" | "key-points" | "general"

gemini-brainstorm

Collaborative brainstorming:

prompt: "How could we implement real-time collaboration?"
claudeThoughts: "I think we should use WebSockets..."
maxRounds: 3 (optional)

gemini-summarize

Summarize content:

content: "Long text to summarize..."
length: "brief" | "moderate" | "detailed"
format: "paragraph" | "bullet-points" | "outline"

gemini-run-code

Let Gemini write and execute Python code:

prompt: "Calculate the first 50 prime numbers and plot them"
data: "optional CSV data to analyze" (optional)

Supports libraries: numpy, pandas, matplotlib, scipy, scikit-learn, tensorflow, and more. Generated charts are saved to the output directory and returned as images.

gemini-search

Real-time web search with citations:

query: "What happened in tech news this week?"
returnCitations: true (default)

Returns grounded responses with inline citations and source URLs.

gemini-structured

Get JSON responses matching a schema:

prompt: "Extract the meeting details from this email..."
schema: '{"type":"object","properties":{"date":{"type":"string"},"attendees":{"type":"array"}}}'
useGoogleSearch: false (optional)

gemini-extract

Convenience tool for common extraction patterns:

text: "Your text to analyze..."
extractType: "entities" | "facts" | "summary" | "keywords" | "sentiment" | "custom"
customFields: "name, date, amount" (for custom extraction)

gemini-youtube

Analyze YouTube videos directly:

url: "https://www.youtube.com/watch?v=..."
question: "What happens at 2:30?"
startTime: "1m30s" (optional, for clipping)
endTime: "5m00s" (optional, for clipping)

gemini-youtube-summary

Quick video summarization:

url: "https://www.youtube.com/watch?v=..."
style: "brief" | "detailed" | "bullet-points" | "chapters"

gemini-analyze-document

Analyze PDFs and documents:

filePath: "/path/to/document.pdf"
question: "Summarize the key findings"
mediaResolution: "low" | "medium" | "high"

gemini-summarize-pdf

Quick PDF summarization:

filePath: "/path/to/document.pdf"
style: "brief" | "detailed" | "outline" | "key-points"

gemini-extract-tables

Extract tables from documents:

filePath: "/path/to/document.pdf"
outputFormat: "markdown" | "csv" | "json"

Workflow: Claude + Gemini

The killer combination for development:

Claude	Gemini
Complex logic	Frontend/UI
Architecture	Visual components
Backend code	Image generation
Integration	React/CSS styling
Reasoning	Creative generation

Example workflow:

Ask Claude to design the backend API
Use gemini-generate-image for UI mockups
Ask Gemini to generate React components via gemini-query
Use multi-turn editing to refine visuals
Let Claude wire everything together

Environment Variables

Variable	Required	Default	Description
`GEMINI_API_KEY`	Yes	-	Your Google Gemini API key
`GEMINI_OUTPUT_DIR`	No	`./gemini-output`	Where to save generated files
`GEMINI_MODEL`	No	-	Override model for init test
`GEMINI_PRO_MODEL`	No	`gemini-3-pro-preview`	Pro model (Gemini 3)
`GEMINI_FLASH_MODEL`	No	`gemini-3-flash-preview`	Flash model (Gemini 3)
`GEMINI_IMAGE_MODEL`	No	`gemini-3-pro-image-preview`	Image model (Nano Banana Pro)
`GEMINI_IMAGE_THINKING_LEVEL`	No	`high`	Default thinking level for image generation (minimal, low, medium, high)
`GEMINI_VIDEO_MODEL`	No	`veo-2.0-generate-001`	Video model
`VERBOSE`	No	`false`	Enable verbose logging
`QUIET`	No	`false`	Minimize logging
`GEMINI_ENABLED_TOOLS`	No	-	Comma-separated list of tool groups to load (e.g., `query,search,image-gen`)
`GEMINI_TOOL_PRESET`	No	-	Preset profile: `minimal`, `text`, `image`, `research`, `media`, `full`

Tool Configuration

By default, all 37 tools are loaded. To reduce context usage, configure which tools to load:

Available Presets

Preset	Tool Groups
`minimal`	query, brainstorm
`text`	query, brainstorm, analyze, summarize, structured
`image`	query, image-gen, image-edit, image-analyze
`research`	query, search, deep-research, url-context, document
`media`	query, image-gen, image-edit, image-analyze, video-gen, youtube, speech
`full`	All 18 tool groups (default)

Using Presets

# Minimal - query and brainstorm
GEMINI_TOOL_PRESET=minimal

# Text processing
GEMINI_TOOL_PRESET=text  # query, brainstorm, analyze, summarize, structured

# Image workflows
GEMINI_TOOL_PRESET=image  # query, image-gen, image-edit, image-analyze

# Research workflows
GEMINI_TOOL_PRESET=research  # query, search, deep-research, url-context, document

Using Explicit Tool Lists

# Only specific tools
GEMINI_ENABLED_TOOLS=query,search,image-gen

Combining Preset + Explicit

# Start with preset, add extras
GEMINI_TOOL_PRESET=minimal
GEMINI_ENABLED_TOOLS=search,image-gen  # Adds to minimal preset

Available Tool Groups

Group	Tools
`query`	gemini-query
`brainstorm`	gemini-brainstorm
`analyze`	gemini-analyze-code, gemini-analyze-text
`summarize`	gemini-summarize
`image-gen`	gemini-generate-image, gemini-image-prompt
`image-edit`	gemini-start-image-edit, gemini-continue-image-edit, gemini-end-image-edit, gemini-list-image-sessions
`video-gen`	gemini-generate-video, gemini-check-video
`code-exec`	gemini-run-code
`search`	gemini-search
`structured`	gemini-structured, gemini-extract
`youtube`	gemini-youtube, gemini-youtube-summary
`document`	gemini-analyze-document, gemini-summarize-pdf, gemini-extract-tables
`url-context`	gemini-analyze-url, gemini-compare-urls, gemini-extract-from-url
`cache`	gemini-create-cache, gemini-query-cache, gemini-list-caches, gemini-delete-cache
`speech`	gemini-speak, gemini-dialogue, gemini-list-voices
`token-count`	gemini-count-tokens
`deep-research`	gemini-deep-research, gemini-check-research, gemini-research-followup
`image-analyze`	gemini-analyze-image

Manual Installation

Global Install

# Using npm
npm install -g @rlabs-inc/gemini-mcp

# Using bun
bun install -g @rlabs-inc/gemini-mcp

Claude Code Configuration

{
  "gemini": {
    "command": "npx",
    "args": ["-y", "@rlabs-inc/gemini-mcp"],
    "env": {
      "GEMINI_API_KEY": "your-api-key",
      "GEMINI_OUTPUT_DIR": "/path/to/save/files"
    }
  }
}

Troubleshooting

Rate Limits (429 Errors)

If you're hitting rate limits on the free tier:

Set GEMINI_MODEL=gemini-3-flash-preview to use Flash for init (higher limits)
Or upgrade to a paid plan

Connection Issues

Verify your API key at Google AI Studio
Check server status: claude mcp list
Try with verbose logging: VERBOSE=true

Image/Video Issues

Ensure your API key has access to image/video generation
Check output directory permissions
Files save to GEMINI_OUTPUT_DIR (default: ./gemini-output)
For 4K images, generation takes longer

Previous Versions

0.7.2

Beautiful CLI with Themes! Use Gemini directly from your terminal:

# Install globally
npm install -g @rlabs-inc/gemini-mcp

# Set your API key once
gcli config set api-key YOUR_KEY

# Generate images, videos, search, research, and more!
gcli image "a cat astronaut" --size 4K
gcli search "latest AI news"
gcli research "quantum computing applications" --wait
gcli speak "Hello world" --voice Puck

5 Beautiful Themes: terminal, neon, ocean, forest, minimal

CLI Commands:

gcli query - Direct Gemini queries with thinking levels
gcli search - Real-time web search with citations
gcli research - Deep research agent
gcli image - Generate images (up to 4K)
gcli video - Generate videos with Veo
gcli speak - Text-to-speech with 30 voices
gcli tokens - Count tokens and estimate costs
gcli config - Manage settings

v0.6.x: Deep Research, Token Counting, TTS, URL analysis, Context Caching v0.5.x: 30+ tools, YouTube analysis, Document analysis v0.4.x: Code execution, Google Search v0.3.x: Thinking levels, Structured output, 4K images v0.2.x: Image/Video generation with Veo

Development

git clone https://github.com/rlabs-inc/gemini-mcp.git
cd gemini-mcp
bun install
bun run build
bun run dev -- --verbose

Scripts

Command	Description
`bun run build`	Build for production
`bun run dev`	Development mode with watch
`bun run typecheck`	Type check without emitting
`bun run format`	Format with Prettier
`bun run lint`	Lint with ESLint

License

MIT License

Made with Claude + Gemini working together

Featured

CodeRabbit

AI writes the code. CodeRabbit catches the slop.

Try For Free →

Make your agent a DeFi expert

Agent, run crypto. Access onchain data & trade routes via 1inch.

Install now →

Make money from your Skills

On Capafy, your Skill runs online 24/7 as an agent product, and you get paid every time someone uses it.

Start earning →

AppSignal

Monitor with ease. Code with confidence.

Start Free Trial →

Vibe Prospecting MCP

Connect Claude to +800M contacts, +150M companies. Find & Enrich leads in chat.

Try For Free →

Context.dev

Integrate web data into your AI product. One API to scrape website & brand data.

Get API Key Now →

Configuration

GEMINI_API_KEY*secret

Your Google Gemini API key (get one at https://aistudio.google.com/apikey)

GEMINI_OUTPUT_DIR

Directory for generated files (images, videos, audio)

MCP Server Gemini

A Model Context Protocol (MCP) server for integrating Google's Gemini 3 models with Claude Code, enabling powerful collaboration between both AI systems. Now with a beautiful CLI!

MCP Registry Support: Now discoverable in the official MCP ecosystem!

Features

Feature	Description
Deep Research Agent	Autonomous multi-step research with web search and citations
Token Counting	Count tokens and estimate costs before API calls
Text-to-Speech	30 unique voices, single speaker or two-speaker dialogues
URL Analysis	Analyze, compare, and extract data from web pages
Context Caching	Cache large documents for efficient repeated queries
YouTube Analysis	Analyze videos by URL with timestamp clipping
Document Analysis	PDFs, DOCX, spreadsheets with table extraction
4K Image Generation	Generate images up to 4K with 10 aspect ratios
Multi-Turn Image Editing	Iteratively refine images through conversation
Video Generation	Create videos with Veo 2.0 (async with polling)
Code Execution	Gemini writes and runs Python code (pandas, numpy, matplotlib)
Google Search	Real-time web information with inline citations
Structured Output	JSON responses with schema validation
Data Extraction	Extract entities, facts, sentiment from text
Thinking Levels	Control reasoning depth (minimal/low/medium/high)
Direct Query	Send prompts to Gemini 3 Pro/Flash models
Brainstorming	Claude + Gemini collaborative problem-solving
Code Analysis	Analyze code for quality, security, performance
Summarization	Summarize content at different detail levels

Quick Installation

MCP Server for Claude Code

# Using npm (Recommended)
claude mcp add gemini -s user -- env GEMINI_API_KEY=YOUR_KEY npx -y @rlabs-inc/gemini-mcp

# Using bun
claude mcp add gemini -s user -- env GEMINI_API_KEY=YOUR_KEY bunx @rlabs-inc/gemini-mcp

CLI (Global Install)

# Install globally
npm install -g @rlabs-inc/gemini-mcp

# Set your API key once (stored securely)
gcli config set api-key YOUR_KEY

# Now use any command!
gcli search "latest news"
glci image "sunset over mountains" --ratio 16:9

Get your API key: Visit Google AI Studio - it's free and takes seconds!

Installation Options

# With verbose logging
claude mcp add gemini -s user -- env GEMINI_API_KEY=YOUR_KEY VERBOSE=true bunx -y @rlabs-inc/gemini-mcp

# With custom output directory for generated images/videos
claude mcp add gemini -s user -- env GEMINI_API_KEY=YOUR_KEY GEMINI_OUTPUT_DIR=/path/to/output bunx -y @rlabs-inc/gemini-mcp

Available Tools

gemini-query

Direct queries to Gemini with thinking level control:

prompt: "Explain quantum entanglement"
model: "pro" or "flash"
thinkingLevel: "low" | "medium" | "high" (optional)

low: Fast responses, minimal reasoning
medium: Balanced (Flash only)
high: Deep reasoning for complex tasks (default)

gemini-generate-image

Generate images with Nano Banana Pro (Claude can SEE them!):

prompt: "a futuristic city at sunset"
style: "cyberpunk" (optional)
aspectRatio: "16:9" (1:1, 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9)
imageSize: "2K" (1K, 2K, 4K)
useGoogleSearch: false (ground in real-world info)
thinkingLevel: "high" (optional - minimal, low, medium, high)
personGeneration: "ALLOW_ALL" (optional - ALLOW_ALL, ALLOW_ADULT, ALLOW_NONE)
seed: 42 (optional - for reproducible results)

gemini-start-image-edit

Start a multi-turn image editing session:

prompt: "a cozy cabin in the mountains"
aspectRatio: "16:9"
imageSize: "2K"
useGoogleSearch: false
thinkingLevel: "high" (optional - minimal, low, medium, high)
personGeneration: "ALLOW_ALL" (optional - ALLOW_ALL, ALLOW_ADULT, ALLOW_NONE)
seed: 42 (optional - for reproducible results)

Returns a session ID for iterative editing.

gemini-continue-image-edit

Continue refining an image:

sessionId: "edit-123456789"
prompt: "add snow on the roof and make it nighttime"

gemini-end-image-edit

Close an editing session:

sessionId: "edit-123456789"

gemini-list-image-sessions

List all active editing sessions.

gemini-generate-video

Generate videos using Veo:

prompt: "a cat playing piano"
aspectRatio: "16:9" (optional)
negativePrompt: "blurry, text" (optional)

Video generation is async (takes 1-5 minutes). Use gemini-check-video to poll.

gemini-check-video

Check video generation status and download when complete:

operationId: "operations/xxx-xxx-xxx"

gemini-analyze-code

Analyze code for issues:

code: "function foo() { ... }"
language: "typescript" (optional)
focus: "quality" | "security" | "performance" | "bugs" | "general"

gemini-analyze-text

Analyze text content:

text: "Your text here..."
type: "sentiment" | "summary" | "entities" | "key-points" | "general"

gemini-brainstorm

Collaborative brainstorming:

prompt: "How could we implement real-time collaboration?"
claudeThoughts: "I think we should use WebSockets..."
maxRounds: 3 (optional)

gemini-summarize

Summarize content:

content: "Long text to summarize..."
length: "brief" | "moderate" | "detailed"
format: "paragraph" | "bullet-points" | "outline"

gemini-run-code

Let Gemini write and execute Python code:

prompt: "Calculate the first 50 prime numbers and plot them"
data: "optional CSV data to analyze" (optional)

Supports libraries: numpy, pandas, matplotlib, scipy, scikit-learn, tensorflow, and more. Generated charts are saved to the output directory and returned as images.

gemini-search

Real-time web search with citations:

query: "What happened in tech news this week?"
returnCitations: true (default)

Returns grounded responses with inline citations and source URLs.

gemini-structured

Get JSON responses matching a schema:

prompt: "Extract the meeting details from this email..."
schema: '{"type":"object","properties":{"date":{"type":"string"},"attendees":{"type":"array"}}}'
useGoogleSearch: false (optional)

gemini-extract

Convenience tool for common extraction patterns:

text: "Your text to analyze..."
extractType: "entities" | "facts" | "summary" | "keywords" | "sentiment" | "custom"
customFields: "name, date, amount" (for custom extraction)

gemini-youtube

Analyze YouTube videos directly:

url: "https://www.youtube.com/watch?v=..."
question: "What happens at 2:30?"
startTime: "1m30s" (optional, for clipping)
endTime: "5m00s" (optional, for clipping)

gemini-youtube-summary

Quick video summarization:

url: "https://www.youtube.com/watch?v=..."
style: "brief" | "detailed" | "bullet-points" | "chapters"

gemini-analyze-document

Analyze PDFs and documents:

filePath: "/path/to/document.pdf"
question: "Summarize the key findings"
mediaResolution: "low" | "medium" | "high"

gemini-summarize-pdf

Quick PDF summarization:

filePath: "/path/to/document.pdf"
style: "brief" | "detailed" | "outline" | "key-points"

gemini-extract-tables

Extract tables from documents:

filePath: "/path/to/document.pdf"
outputFormat: "markdown" | "csv" | "json"

Workflow: Claude + Gemini

The killer combination for development:

Claude	Gemini
Complex logic	Frontend/UI
Architecture	Visual components
Backend code	Image generation
Integration	React/CSS styling
Reasoning	Creative generation

Example workflow:

Ask Claude to design the backend API
Use gemini-generate-image for UI mockups
Ask Gemini to generate React components via gemini-query
Use multi-turn editing to refine visuals
Let Claude wire everything together

Environment Variables

Variable	Required	Default	Description
`GEMINI_API_KEY`	Yes	-	Your Google Gemini API key
`GEMINI_OUTPUT_DIR`	No	`./gemini-output`	Where to save generated files
`GEMINI_MODEL`	No	-	Override model for init test
`GEMINI_PRO_MODEL`	No	`gemini-3-pro-preview`	Pro model (Gemini 3)
`GEMINI_FLASH_MODEL`	No	`gemini-3-flash-preview`	Flash model (Gemini 3)
`GEMINI_IMAGE_MODEL`	No	`gemini-3-pro-image-preview`	Image model (Nano Banana Pro)
`GEMINI_IMAGE_THINKING_LEVEL`	No	`high`	Default thinking level for image generation (minimal, low, medium, high)
`GEMINI_VIDEO_MODEL`	No	`veo-2.0-generate-001`	Video model
`VERBOSE`	No	`false`	Enable verbose logging
`QUIET`	No	`false`	Minimize logging
`GEMINI_ENABLED_TOOLS`	No	-	Comma-separated list of tool groups to load (e.g., `query,search,image-gen`)
`GEMINI_TOOL_PRESET`	No	-	Preset profile: `minimal`, `text`, `image`, `research`, `media`, `full`

Tool Configuration

By default, all 37 tools are loaded. To reduce context usage, configure which tools to load:

Available Presets

Preset	Tool Groups
`minimal`	query, brainstorm
`text`	query, brainstorm, analyze, summarize, structured
`image`	query, image-gen, image-edit, image-analyze
`research`	query, search, deep-research, url-context, document
`media`	query, image-gen, image-edit, image-analyze, video-gen, youtube, speech
`full`	All 18 tool groups (default)

Using Presets

# Minimal - query and brainstorm
GEMINI_TOOL_PRESET=minimal

# Text processing
GEMINI_TOOL_PRESET=text  # query, brainstorm, analyze, summarize, structured

# Image workflows
GEMINI_TOOL_PRESET=image  # query, image-gen, image-edit, image-analyze

# Research workflows
GEMINI_TOOL_PRESET=research  # query, search, deep-research, url-context, document

Using Explicit Tool Lists

# Only specific tools
GEMINI_ENABLED_TOOLS=query,search,image-gen

Combining Preset + Explicit

# Start with preset, add extras
GEMINI_TOOL_PRESET=minimal
GEMINI_ENABLED_TOOLS=search,image-gen  # Adds to minimal preset

Available Tool Groups

Group	Tools
`query`	gemini-query
`brainstorm`	gemini-brainstorm
`analyze`	gemini-analyze-code, gemini-analyze-text
`summarize`	gemini-summarize
`image-gen`	gemini-generate-image, gemini-image-prompt
`image-edit`	gemini-start-image-edit, gemini-continue-image-edit, gemini-end-image-edit, gemini-list-image-sessions
`video-gen`	gemini-generate-video, gemini-check-video
`code-exec`	gemini-run-code
`search`	gemini-search
`structured`	gemini-structured, gemini-extract
`youtube`	gemini-youtube, gemini-youtube-summary
`document`	gemini-analyze-document, gemini-summarize-pdf, gemini-extract-tables
`url-context`	gemini-analyze-url, gemini-compare-urls, gemini-extract-from-url
`cache`	gemini-create-cache, gemini-query-cache, gemini-list-caches, gemini-delete-cache
`speech`	gemini-speak, gemini-dialogue, gemini-list-voices
`token-count`	gemini-count-tokens
`deep-research`	gemini-deep-research, gemini-check-research, gemini-research-followup
`image-analyze`	gemini-analyze-image

Manual Installation

Global Install

# Using npm
npm install -g @rlabs-inc/gemini-mcp

# Using bun
bun install -g @rlabs-inc/gemini-mcp

Claude Code Configuration

{
  "gemini": {
    "command": "npx",
    "args": ["-y", "@rlabs-inc/gemini-mcp"],
    "env": {
      "GEMINI_API_KEY": "your-api-key",
      "GEMINI_OUTPUT_DIR": "/path/to/save/files"
    }
  }
}

Troubleshooting

Rate Limits (429 Errors)

If you're hitting rate limits on the free tier:

Set GEMINI_MODEL=gemini-3-flash-preview to use Flash for init (higher limits)
Or upgrade to a paid plan

Connection Issues

Verify your API key at Google AI Studio
Check server status: claude mcp list
Try with verbose logging: VERBOSE=true

Image/Video Issues

Ensure your API key has access to image/video generation
Check output directory permissions
Files save to GEMINI_OUTPUT_DIR (default: ./gemini-output)
For 4K images, generation takes longer

Previous Versions

0.7.2

Beautiful CLI with Themes! Use Gemini directly from your terminal:

# Install globally
npm install -g @rlabs-inc/gemini-mcp

# Set your API key once
gcli config set api-key YOUR_KEY

# Generate images, videos, search, research, and more!
gcli image "a cat astronaut" --size 4K
gcli search "latest AI news"
gcli research "quantum computing applications" --wait
gcli speak "Hello world" --voice Puck

5 Beautiful Themes: terminal, neon, ocean, forest, minimal

CLI Commands:

gcli query - Direct Gemini queries with thinking levels
gcli search - Real-time web search with citations
gcli research - Deep research agent
gcli image - Generate images (up to 4K)
gcli video - Generate videos with Veo
gcli speak - Text-to-speech with 30 voices
gcli tokens - Count tokens and estimate costs
gcli config - Manage settings

Development

git clone https://github.com/rlabs-inc/gemini-mcp.git
cd gemini-mcp
bun install
bun run build
bun run dev -- --verbose

Scripts

Command	Description
`bun run build`	Build for production
`bun run dev`	Development mode with watch
`bun run typecheck`	Type check without emitting
`bun run format`	Format with Prettier
`bun run lint`	Lint with ESLint

License

MIT License

Made with Claude + Gemini working together