
Search PubChem compounds, properties, safety data, bioactivity, and cross-references.
Search the PubChem chemical database for compounds, properties, safety data, bioactivity, cross-references, and entity summaries via MCP. STDIO or Streamable HTTP.
Public Hosted Server: https://pubchem.caseyjhand.com/mcp
Chemical compound and bioassay data from PubChem's PUG REST and PUG View APIs. Search compounds by identifier, formula, or structure; fetch physicochemical properties, safety data, bioactivity, interactions, cross-references, and 3D structures; find bioassays by biological target. Runs as a stdio process, a local Streamable HTTP server, or the public hosted endpoint above.
| Tool | Description |
|---|---|
pubchem_search_compounds | Search for compounds by name, SMILES, InChIKey, formula, substructure, superstructure, or 2D similarity. |
pubchem_get_compound_details | Get physicochemical properties, descriptions, synonyms, drug-likeness, and classification for compounds by CID. |
pubchem_get_compound_image | Fetch a 2D structure diagram (PNG) for a compound by CID. |
pubchem_get_compound_3d_structure | Fetch a 3D conformer (atomic coordinates and bonds) for a compound by CID, as parsed JSON or raw SDF. |
pubchem_get_compound_xrefs | Get external database cross-references (PubMed, patents, genes, proteins, etc.). |
pubchem_get_compound_safety | Get GHS hazard classification and safety data for one or more compounds by CID (batch). |
pubchem_get_bioactivity | Get a compound's bioactivity profile: assay results, targets, and activity values; filter by outcome or molecular target. |
pubchem_get_compound_interactions | Get drug-drug, drug-food, and chemical-target interactions for a compound by CID. |
pubchem_search_assays | Find bioassays by biological target (gene symbol, protein, Gene ID, UniProt accession). |
pubchem_get_summary | Get summaries for PubChem entities: assays, genes, proteins, taxonomy. |
Compound and assay records are also exposed as URI-templated resources, backed by the same client methods as the tools; many MCP clients are tool-only and never surface resources.
| Resource | Description |
|---|---|
pubchem://compound/{cid} | Core physicochemical properties (JSON). |
pubchem://compound/{cid}/safety | GHS hazard classification (JSON). |
pubchem://compound/{cid}/image | 2D structure diagram (PNG). |
pubchem://compound/{cid}/xrefs | External cross-references (JSON). |
pubchem://compound/{cid}/bioactivity | Bioassay activity profile (JSON). |
pubchem://assay/{aid} | BioAssay summary (JSON). |
pubchem_search_compounds toolallowOtherElements), substructure/superstructure containment, or 2D Tanimoto similarity (threshold 70-100, default 90)identifierType + identifiers; formula: formula; substructure/superstructure/similarity: query + queryType — and a missing or blank one is rejected before the upstream calloffset pages to a ceiling of 10,000 — identifier lookups resolve every match up front so paging is free, while formula/structure/similarity searches cost more upstream per deep pageproperties hydration avoids a follow-up pubchem_get_compound_details callunresolvedIdentifiers for inputs that resolved to no CID — no PubChem match, or a SMILES PubChem cannot interpret — while the rest of the batch still resolves, plus notices when multiple inputs collide on one CID* wildcard atom, a CID with no record) fails fast with a search_query_rejected hint naming what to fixtotalFound when the full match set was observed, or a totalFoundAtLeast floor when a bounded upstream search saturatedpubchem_get_compound_details tooldescriptionOffset/maxDescriptions (default 3, up to 20) — fetched only for the first 10 CIDs in the batch, remaining CIDs listed in skippedCidssynonymOffset/maxSynonyms (default 20, up to 100)found: false distinguishes a nonexistent CID from a real compound PubChem simply has no data forpubchem_get_compound_image toolsize is "small" (100x100) or "large" (300x300, default)cid_not_found error when PubChem has no record for the CIDpubchem_get_compound_3d_structure toolformat="json" (default) returns parsed atoms (element + x/y/z) and bonds, format="sdf" returns the raw V2000 SDF textmaxAtoms/maxBonds cap the JSON preview (default 200 each); atomCount/bondCount always report the full totals, with any capping disclosed via enrichmentincludeRawSdf bypasses the default 500-line cap on the raw SDF textincludeAlternateConformerIds lists conformer IDs beyond the defaultno_3d_structure error when PubChem has no computed 3D coordinates (large molecules, mixtures, some salts)pubchem_get_compound_xrefs toolxrefTypes — string IDs (RegistryID, RN for CAS numbers, PatentID) and numeric IDs (PubMedID, GeneID, ProteinGI, TaxonomyID)maxPerType up to 500 (default 50), with the same offset applied across every requested typetotalAvailable and truncated flagpubchem_get_compound_safety toolstatus: ok, no_ghs_data (compound exists, no deposited classification), or cid_not_found (no PubChem record at all) — kept distinct so a bad CID never reads as "no hazards on file"decoded flag — false for codes needing label-specific fill text or outside the decoder table; the code itself is still authoritativepubchem_get_bioactivity tooloutcomeFilter (active/inactive/all, default all) and/or targetGeneId/targetAccessionoffset reaches the resttotalAssays/activeCount/inactiveCount for the whole compound, plus filteredCount/returnedCount for the current pagepubchem_get_compound_interactions toolkinds — drug-drug (DrugBank), drug-food, target (binding/activity from BindingDB, ChEMBL, and others); default ["drug-drug"]maxEntries per kind per page (1-50, default 10); offset counts source records rather than returned entries, capped at 2,147,483,646paging[] reports per-kind totalRecords/nextOffset/truncated; the top-level nextOffset is populated only when exactly one requested kind still has records leftfailedKinds without failing the kinds that succeededpubchem_search_assays tooltargetType: genesymbol/proteinname (text), geneid (NCBI Gene ID), proteinaccession (UniProt)offset pages to the total foundtargetQuery and a non-numeric geneid query before the upstream calltotalFound across all pages and distinguishes "no match" from "offset past the end"pubchem_get_summary toolentityType: assay (AID), gene (NCBI Gene ID), protein (UniProt accession), or taxonomy (Tax ID); up to 10 identifiers per callfound flag; populated fields depend on entityType (taxonomy includes an ordered lineage, gene includes symbol/taxonomy)entityType expectspubchem://compound/{cid} resourcepubchem_get_compound_details), as application/jsonpubchem_get_compound_details to select specific properties or add descriptions, synonyms, drug-likeness, and classificationpubchem://compound/{cid}/safety resourceapplication/jsonstatus (ok/no_ghs_data/cid_not_found) is the only signal distinguishing a bad CID from a compound with no deposited classification — a resource read has no notice surfacepubchem://compound/{cid}/image resourcepubchem_get_compound_image for the 100x100 size optionpubchem://compound/{cid}/xrefs resourceRN (CAS), RegistryID, PubMedID — up to 25 IDs per type, as application/jsonpubchem_get_compound_xrefs for the full set of xref types, a higher per-type cap, and offset pagingpubchem://compound/{cid}/bioactivity resourceapplication/json, plus totalAssays/activeCount for the whole compoundpubchem_get_bioactivity to filter by outcome or target, raise the cap, or page with offsetpubchem://assay/{aid} resourceapplication/json — name, description, source, protocol, substance countsBuilt on @cyanheads/mcp-ts-core: stdio and Streamable HTTP transports, pluggable auth (none / jwt / oauth), swappable storage (in-memory, filesystem, Supabase, Cloudflare KV/R2/D1), structured logging with optional OpenTelemetry tracing.
PubChem-specific:
RequestCancelledAgent-friendly output:
status (ok / no_ghs_data / cid_not_found) and found flags let callers branch on data instead of matching an error stringunresolvedIdentifiers, skippedCids, and failedKinds rather than failing the whole calltruncated, shown/cap, nextOffset) on every capped list, plus a totalFoundAtLeast floor in place of a count when an upstream search saturatesreason (e.g. cid_not_found, missing_identifier_args, invalid_cid_query) with actionable recovery text, not generic messagesA public instance is available at https://pubchem.caseyjhand.com/mcp — no installation required. Point any MCP client at it via Streamable HTTP:
{
"mcpServers": {
"pubchem-mcp-server": {
"type": "streamable-http",
"url": "https://pubchem.caseyjhand.com/mcp"
}
}
}
Add the following to your MCP client configuration file.
{
"mcpServers": {
"pubchem-mcp-server": {
"type": "stdio",
"command": "bunx",
"args": ["@cyanheads/pubchem-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio"
}
}
}
}
Or with npx (no Bun required):
{
"mcpServers": {
"pubchem-mcp-server": {
"type": "stdio",
"command": "npx",
"args": ["-y", "@cyanheads/pubchem-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio"
}
}
}
}
Or with Docker:
{
"mcpServers": {
"pubchem-mcp-server": {
"type": "stdio",
"command": "docker",
"args": ["run", "-i", "--rm", "-e", "MCP_TRANSPORT_TYPE=stdio", "ghcr.io/cyanheads/pubchem-mcp-server:latest"]
}
}
}
For Streamable HTTP, set the transport and start the server:
MCP_TRANSPORT_TYPE=http MCP_HTTP_PORT=3010 bun run start:http
# Server listens at http://localhost:3010/mcp
git clone https://github.com/cyanheads/pubchem-mcp-server.git
cd pubchem-mcp-server
bun install
cp .env.example .env
# edit .env to override transport, session mode, storage, or logging defaults
| Variable | Description | Default |
|---|---|---|
MCP_TRANSPORT_TYPE | Transport: stdio or http. | stdio |
MCP_HTTP_PORT | Port for HTTP server. | 3010 |
MCP_HTTP_HOST | Host for HTTP server. | 127.0.0.1 |
MCP_SESSION_MODE | stateless, stateful, or auto. PubChem needs no multi-round-trip input, so the server declares stateless; the example and Docker set it to match. | stateless |
MCP_AUTH_MODE | Auth mode: none, jwt, or oauth. | none |
MCP_LOG_LEVEL | Log level (RFC 5424). | info |
STORAGE_PROVIDER_TYPE | Storage backend. | in-memory |
OTEL_ENABLED | Enable OpenTelemetry. | false |
See .env.example for the full list of optional overrides.
Build and run:
# One-time build
bun run rebuild
# Run the built server
bun run start:stdio
# or
bun run start:http
Run checks and tests:
bun run devcheck # Lint, format, typecheck, security
bun run test # Vitest test suite
bun run lint:mcp # Validate MCP definitions against spec
docker build -t pubchem-mcp-server .
docker run --rm -p 3010:3010 pubchem-mcp-server
The Dockerfile defaults to HTTP transport, stateless session mode, and logs to /var/log/pubchem-mcp-server. OpenTelemetry peer dependencies are installed by default — build with --build-arg OTEL_ENABLED=false to omit them.
| Directory | Purpose |
|---|---|
src/index.ts | createApp() entry point — registers tools/resources and inits the PubChem client. |
src/mcp-server/tools/definitions/ | Tool definitions (*.tool.ts). |
src/mcp-server/resources/definitions/ | Resource definitions (*.resource.ts). |
src/services/pubchem/ | PubChem API client — rate limiting, retry, and response/SDF parsing. |
scripts/ | Build, clean, devcheck, and tree generation scripts. |
tests/ | Unit and integration tests. |
See CLAUDE.md for development guidelines and architectural rules. The short version:
try/catch in tool logicctx.log for request-scoped loggingindex.ts barrel filesIssues are welcome. Run checks before submitting:
bun run devcheck
bun run test
Apache-2.0 — see LICENSE for details.