This server wraps the GoldenMatch entity resolution toolkit, which finds duplicate records across messy datasets with a 97.2% F1 score out of the box. You get access to deduplication, clustering, and the broader Golden Suite pipeline (InferMap for schema alignment, GoldenCheck for profiling, GoldenFlow for standardization). The MCP layer exposes 36+ tools including auto_configure for adaptive tuning and controller_telemetry for inspecting clustering decisions. Useful when you need to clean customer lists, merge data sources, or resolve entities across organizations without writing custom fuzzy matching logic. The zero-config defaults work immediately, and a learning memory system stops asking for the same correction twice across runs.
claude mcp add --transport http goldenmatch https://goldenmatch-mcp-production.up.railway.app/mcp/Run in your terminal. Add --scope user to make it available in every project.
Review the command, arguments, and environment values before installing — MCP servers run with your local permissions.
Verified live against the running server on Jun 10, 2026.
analyze_dataProfile data, detect domain, recommend ER strategy1 paramsProfile data, detect domain, recommend ER strategy
file_path*stringauto_configureRun AutoConfigController on a CSV; return the committed GoldenMatchConfig (incl. negative_evidence / Path Y when chosen) plus telemetry — stop_reason, health, decision trace, indicator column priors. Programmatic equivalent of `goldenmatch autoconfig`.2 paramsRun AutoConfigController on a CSV; return the committed GoldenMatchConfig (incl. negative_evidence / Path Y when chosen) plus telemetry — stop_reason, health, decision trace, indicator column priors. Programmatic equivalent of `goldenmatch autoconfig`.
file_path*stringconstraintsobjectcontroller_telemetryReturn the AutoConfigController telemetry from the most recent `auto_configure` or `agent_deduplicate` call in this MCP session. Same JSON shape as the web /api/v1/controller/telemetry endpoint.Return the AutoConfigController telemetry from the most recent `auto_configure` or `agent_deduplicate` call in this MCP session. Same JSON shape as the web /api/v1/controller/telemetry endpoint.
No parameters — call it with no arguments.
agent_deduplicateRun full ER pipeline with confidence gating and reasoning2 paramsRun full ER pipeline with confidence gating and reasoning
configobjectfile_path*stringagent_match_sourcesMatch two files with intelligent strategy selection3 paramsMatch two files with intelligent strategy selection
configobjectfile_a*stringfile_b*stringagent_explain_pairNatural language explanation for a record pair4 paramsNatural language explanation for a record pair
exactarrayfuzzyobjectrecord_a*objectrecord_b*objectagent_explain_clusterExplain why records are in the same cluster1 paramsExplain why records are in the same cluster
cluster_id*integeragent_review_queueGet borderline pairs awaiting approval1 paramsGet borderline pairs awaiting approval
job_name*stringagent_approve_rejectApprove or reject a review queue pair6 paramsApprove or reject a review queue pair
id_a*integerid_b*integerreasonstringdecision*stringjob_name*stringdecided_by*stringagent_compare_strategiesCompare ER strategies on your data2 paramsCompare ER strategies on your data
file_path*stringground_truthstringsuggest_pprlCheck if data needs privacy-preserving matching1 paramsCheck if data needs privacy-preserving matching
file_path*stringscan_qualityRun GoldenCheck data quality scan on a CSV file. Returns issues found (encoding errors, Unicode problems, format violations) without applying fixes. Requires goldencheck: pip install goldenmatch[quality]2 paramsRun GoldenCheck data quality scan on a CSV file. Returns issues found (encoding errors, Unicode problems, format violations) without applying fixes. Requires goldencheck: pip install goldenmatch[quality]
domainstringfile_path*stringfix_qualityRun GoldenCheck scan and apply fixes to a CSV file. Returns the fixed data summary and a manifest of all fixes applied. Requires goldencheck: pip install goldenmatch[quality]4 paramsRun GoldenCheck scan and apply fixes to a CSV file. Returns the fixed data summary and a manifest of all fixes applied. Requires goldencheck: pip install goldenmatch[quality]
domainstringfix_modestringsafe · moderatedefault: safefile_path*stringoutput_pathstringrun_transformsRun GoldenFlow data transforms on a CSV file. Normalizes phone numbers (E.164), dates (ISO), categorical spelling, and Unicode issues. Returns a manifest of transforms applied. Requires goldenflow: pip install goldenmatch[transform]2 paramsRun GoldenFlow data transforms on a CSV file. Normalizes phone numbers (E.164), dates (ISO), categorical spelling, and Unicode issues. Returns a manifest of transforms applied. Requires goldenflow: pip install goldenmatch[transform]
file_path*stringoutput_pathstringlist_correctionsList stored Learning Memory corrections, optionally filtered by dataset. Returns id_a, id_b, decision, source, trust, reason, matchkey_name, dataset, original_score, created_at.2 paramsList stored Learning Memory corrections, optionally filtered by dataset. Returns id_a, id_b, decision, source, trust, reason, matchkey_name, dataset, original_score, created_at.
pathstringdatasetstringadd_correctionAdd a pair correction to Learning Memory. Source is set to 'agent' with trust=0.5 (lower than human steward decisions which are 1.0). Pair (id_a, id_b) is canonicalized to (min, max) before storage.7 paramsAdd a pair correction to Learning Memory. Source is set to 'agent' with trust=0.5 (lower than human steward decisions which are 1.0). Pair (id_a, id_b) is canonicalized to (min, max) before storage.
id_a*integerid_b*integerpathstringreasonstringdataset*stringdecision*stringapprove · rejectmatchkey_namestringlearn_thresholdsForce a MemoryLearner pass over accumulated corrections. Returns the list of LearnedAdjustments produced (matchkey_name, threshold, sample_size, learned_at). Requires >= 10 corrections per matchkey before threshold tuning fires; otherwise returns an empty list.2 paramsForce a MemoryLearner pass over accumulated corrections. Returns the list of LearnedAdjustments produced (matchkey_name, threshold, sample_size, learned_at). Requires >= 10 corrections per matchkey before threshold tuning fires; otherwise returns an empty list.
pathstringmatchkey_namestringmemory_statsReturn Learning Memory status: total correction count, last learn time, and current learned adjustments. Cheap; safe for status checks.1 paramsReturn Learning Memory status: total correction count, last learn time, and current learned adjustments. Cheap; safe for status checks.
pathstringmemory_exportReturn all corrections as a list of dicts (CSV-shaped). Caller is responsible for writing the file. Optionally filter by dataset.2 paramsReturn all corrections as a list of dicts (CSV-shaped). Caller is responsible for writing the file. Optionally filter by dataset.
pathstringdatasetstringidentity_resolveResolve a record_id to its durable identity. Returns the full identity view (members, evidence edges, recent events) or null when no identity exists for that record.2 paramsResolve a record_id to its durable identity. Returns the full identity view (members, evidence edges, recent events) or null when no identity exists for that record.
pathstringrecord_id*stringidentity_listList identities, optionally filtered by dataset/status.5 paramsList identities, optionally filtered by dataset/status.
pathstringlimitintegeroffsetintegerstatusstringdatasetstringidentity_historyReturn the temporal event log for an identity.3 paramsReturn the temporal event log for an identity.
pathstringlimitintegerentity_id*stringidentity_conflictsList evidence edges marked `conflicts_with`.2 paramsList evidence edges marked `conflicts_with`.
pathstringdatasetstringidentity_mergeManually merge two identities. All records from `absorb_entity_id` are reassigned to `keep_entity_id`.4 paramsManually merge two identities. All records from `absorb_entity_id` are reassigned to `keep_entity_id`.
pathstringreasonstringkeep_entity_id*stringabsorb_entity_id*stringidentity_splitSplit a subset of records off an identity into a brand-new identity. The original keeps the remaining records.4 paramsSplit a subset of records off an identity into a brand-new identity. The original keeps the remaining records.
pathstringreasonstringentity_id*stringrecord_ids*arrayget_statsGet dataset statistics: record count, cluster count, match rate, cluster sizes.Get dataset statistics: record count, cluster count, match rate, cluster sizes.
No parameters — call it with no arguments.
find_duplicatesFind duplicate matches for a record. Provide field values to search against the loaded dataset.2 paramsFind duplicate matches for a record. Provide field values to search against the loaded dataset.
top_kintegerrecord*objectexplain_matchExplain why two records match or don't match. Shows per-field score breakdown.2 paramsExplain why two records match or don't match. Shows per-field score breakdown.
record_a*objectrecord_b*objectlist_clustersList duplicate clusters found in the dataset. Returns cluster IDs, sizes, and member counts.2 paramsList duplicate clusters found in the dataset. Returns cluster IDs, sizes, and member counts.
limitintegermin_sizeintegerget_clusterGet details of a specific cluster: all member records and their field values.1 paramsGet details of a specific cluster: all member records and their field values.
cluster_id*integerget_golden_recordGet the merged golden (canonical) record for a cluster.1 paramsGet the merged golden (canonical) record for a cluster.
cluster_id*integermatch_recordMatch a single record against the loaded dataset in real-time. Paste a record's fields and instantly see if it matches any existing record. Uses the configured matchkeys, scorers, and thresholds. Example: {"name": "John Smith", "email": "john@test.com", "zip": "10001"}3 paramsMatch a single record against the loaded dataset in real-time. Paste a record's fields and instantly see if it matches any existing record. Uses the configured matchkeys, scorers, and thresholds. Example: {"name": "John Smith", "email": "john@test.com", "zip": "10001"}
top_kintegerrecord*objectthresholdnumberunmerge_recordRemove a record from its cluster. The record becomes a singleton. Remaining cluster members are re-clustered using stored pair scores. Use this to fix bad merges.1 paramsRemove a record from its cluster. The record becomes a singleton. Remaining cluster members are re-clustered using stored pair scores. Use this to fix bad merges.
record_id*integershatter_clusterBreak an entire cluster into individual records. All members become singletons. Use when a cluster is completely wrong.1 paramsBreak an entire cluster into individual records. All members become singletons. Use when a cluster is completely wrong.
cluster_id*integersuggest_configAnalyze bad merges and suggest config changes. Provide examples of incorrect merges (pairs that should NOT have matched) and GoldenMatch will identify which fields/thresholds to tighten. Example: [{"record_a": {...}, "record_b": {...}, "reason": "different people"}]1 paramsAnalyze bad merges and suggest config changes. Provide examples of incorrect merges (pairs that should NOT have matched) and GoldenMatch will identify which fields/thresholds to tighten. Example: [{"record_a": {...}, "record_b": {...}, "reason": "different people"}]
bad_merges*arrayprofile_dataGet data quality profile: column types, null rates, unique counts, sample values.Get data quality profile: column types, null rates, unique counts, sample values.
No parameters — call it with no arguments.
export_resultsExport matching results to a file (CSV or JSON).2 paramsExport matching results to a file (CSV or JSON).
formatstringcsv · jsondefault: csvoutput_path*stringlist_domainsList available domain extraction rulebooks (built-in + user-defined).List available domain extraction rulebooks (built-in + user-defined).
No parameters — call it with no arguments.
create_domainCreate a custom domain extraction rulebook. Define patterns for a specific data domain (medical devices, automotive parts, real estate, etc.).7 paramsCreate a custom domain extraction rulebook. Define patterns for a specific data domain (medical devices, automotive parts, real estate, etc.).
name*stringscopestringlocal · globaldefault: localsignals*arraystop_wordsarraybrand_patternsarrayattribute_patternsobjectidentifier_patternsobjecttest_domainTest a domain extraction rulebook against sample records. Shows what features would be extracted from the loaded data.2 paramsTest a domain extraction rulebook against sample records. Shows what features would be extracted from the loaded data.
domain_name*stringsample_sizeintegerpprl_auto_configAnalyze the loaded dataset and recommend optimal PPRL (privacy-preserving record linkage) configuration. Returns recommended fields, bloom filter parameters, threshold, and explanation.2 paramsAnalyze the loaded dataset and recommend optimal PPRL (privacy-preserving record linkage) configuration. Returns recommended fields, bloom filter parameters, threshold, and explanation.
use_llmbooleansecurity_levelstringstandard · high · paranoiddefault: highpprl_linkRun privacy-preserving record linkage between two parties' data. Computes bloom filters, matches records without sharing raw data. Specify fields, threshold, and security level.5 paramsRun privacy-preserving record linkage between two parties' data. Computes bloom filters, matches records without sharing raw data. Specify fields, threshold, and security level.
fields*arrayfile_a*stringfile_b*stringthresholdnumbersecurity_levelstringstandard · high · paranoiddefault: highSplink-beating entity resolution — Arrow-native, Rust-fast, zero-tuning — feeding a durable identity layer, so messy records from every source become stable golden entities with whole-record, Customer-360 provenance.
Zero-config matching that beats expert-tuned Splink head-to-head on messy customer records, in an Arrow-native, Rust-authoritative engine verified from a laptop CSV to a 100M-row dedupe in 9.2 minutes. The identities it produces live in a transaction-native control plane — stable entity_ids, per-field provenance, merge/split, and a tamper-evident audit log — one call away as a Customer 360. It even owns its primitives: byte-identical, faster-than-rapidfuzz / jellyfish / FAISS Rust kernels, not rented dependencies.
Python · TypeScript · SQL, at 4-decimal parity · native in Postgres + DuckDB · edge WASM · 70+ MCP tools · beats hand-tuned Splink · 100M rows in 9.2 min
Pair drilldown in the web workbench: cluster members, field-level diff, and a one-line NL explanation per pair. pip install goldenmatch[web] then goldenmatch serve-ui <project>. More screenshots →
v3.5.0 — New
datescorer for date fields (#1858).jaro_winklerscores unrelated ISO birthdays 0.80+ (the fixedYYYY-MM-DDshape + shared digit alphabet dominate), so it can't tell a typo from a different person. Thedatescorer compares dates by Damerau-Levenshtein over the canonical digits — a typo scores 0.90, an unrelated date 0.00 — with alevenshteinfallback for non-ISO input. Cross-surface (Python, native kernel, TypeScript), and a preflight check warns when a name-oriented scorer sits on a date field.v3.4.0 — Embeddings are first-class on Fellegi-Sunter matchkeys.
embeddingandrecord_embeddingfield scorers now train (EM) and score end-to-end on the probabilistic path via the vectorized matrix — previously they raisedUnknown scoreron both training and scoring. They are matrix-only, so a matchkey carrying one always runs vectorized, and the TUI now routes FS through the same native/vectorized selector.v3.3.0 — 3.3.0 — negative evidence on Fellegi-Sunter matchkeys.
negative_evidencenow works ontype: probabilisticmatchkeys as EM-learned__ne__dimensions (no labels needed;penalty_bitsas a fixed override), and the Splink migration upgrade pass gains a fan-out lever — a risk-gated NE suggestion plus cluster-guard tuning from your reference clusters.goldenmatch-native0.1.15 scores NE in the Rust kernels (FS_SUPPORTS_NE; older wheels keep the pure-Python fallback automatically).
Most entity-resolution tools hand you clusters and stop. GoldenMatch keeps going: it resolves messy records into a durable golden entity — one per real-world customer — that survives re-runs, carries provenance on every field, and answers "who is this, and where did each value come from?" in a single call.
entity_id (UUIDv7) that persists across runs as new data arrives — records are absorbed, entities merge or split, but the id an entity earns is the id downstream systems can rely on. Run-local cluster numbers reshuffle on every run; these don't.customer_360(entity_id) composes it into one read — golden record, per-field provenance, every linked source record, the event timeline, and the entity's relationship neighborhood:
// customer_360("018f...c2a1") — trimmed
{
"entity_id": "018f2b7e-...-c2a1", "confidence": 0.97, "record_count": 3,
"sources": ["salesforce", "billing", "support"],
"golden_record": { "name": "Ada Lovelace", "email": "ada@analytical.io", "phone": "+1-555-0100" },
"field_provenance": [
{ "field": "email", "value": "ada@analytical.io",
"winning_source": "billing", "winning_record_id": "billing:8821",
"conflicting_values": [ { "value": "ada@ada.dev", "source": "salesforce" } ] },
{ "field": "phone", "value": "+1-555-0100", "winning_source": "salesforce" }
],
"timeline": [ { "kind": "created", "actor": "pipeline", "recorded_at": "2026-07-30T..." },
{ "kind": "absorbed_record", "reason": "matched billing:8821" } ],
"relationships": [ { "other_entity_id": "018f...9d0e", "kind": "shares_address" } ]
}
What ships today vs. what's emerging. The identity spine is production-grade and in
main: stableentity_ids, per-field provenance, survivorship, merge/split, the append-only log + audit chain, cross-channel stitching, the relationship overlay, and incremental resolution against a persisted index (a new record resolves without a full re-run). Thecustomer_360()serving view above and the source-registry layer that keeps it fresh from live systems are the newer, actively-landing pieces (the source connectors — Snowflake, BigQuery, Salesforce, HubSpot — ship today; the registry that wires them into the spine is emerging) — see the Customer 360 design + ADR. We label the seam rather than blur it.
The golden entity lives in the control plane; the matching that builds it runs in the compute engine. That split is the next section.
The golden entity above is produced by two engines that optimize for genuinely different things — and keeping them distinct is the architecture, not an implementation detail (ADR 0047).
flowchart LR
src([source records])
e360([golden entities · Customer 360])
subgraph compute ["Identity Compute Engine — Arrow-native, Rust-authoritative"]
match[block · score · cluster]
end
subgraph control ["Identity Control Plane — transaction-native state machine"]
spine[stable ids · survivorship · merge/split · provenance · audit]
end
src --> compute -->|resolution batch + evidence| control --> e360
control -.->|persisted index| compute
| Identity Compute Engine | Identity Control Plane | |
|---|---|---|
| Shape | Arrow at bulk boundaries, Rust-authoritative kernels | Transaction-native state machine (SQLite default · Postgres) |
| Job | Block, score, cluster — throughput, vectorized, deterministic per run | Stable ids, survivorship, merge/split, provenance, append-only audit |
| State | Stateless per call; measurement-driven kernelization | Durable, transactional, replayable, auditable |
| Backends | DataFusion · Ray · Sail are replaceable execution backends, none synonymous with GoldenMatch | Storage backends conform to one externally-observable semantics |
Many surfaces, one answer. The same capabilities reach Python, edge-safe TypeScript (with an opt-in WASM backend running the same Rust kernels), SQL inside PostgreSQL and DuckDB, and MCP / REST / A2A — governed by specification + conformance, not copy-paste. There is one authoritative owner per capability; pure-Python / standalone-TS paths are classified, conformance-tested fallbacks. Where a boundary can't cross byte-for-byte, we measure and label it rather than claim parity.
Why a platform engineer should care:
The identity layer is only as good as the matching underneath it — and the matching starts at zero config. dedupe_df(df) runs with no rules and no training data: it profiles the data, picks a defensible configuration, and returns golden records immediately. The config it chose comes back on result.config — inspectable, diffable, versionable. Never a black box.
historical_50k pairwise F1 0.827 vs 0.757, cluster B³ 0.862 vs 0.788, one shared evaluator, reproducible bake-off. Fuzzy, exact, probabilistic (Fellegi-Sunter), and LLM scorers, with EM-trained weights and calibrated scores.result.suggestions. Each is kept only if it doesn't worsen a health proxy — so a suggestion never makes results worse. dedupe_df(df, heal=True) applies and re-runs in one call. You close the gap to expert-tuned without being the expert.Runs on unstructured input, too: extract records from PDFs and images, then resolve them like any other source (
pip install goldenmatch[documents]).
The engine and the identity layer reach your stack through the surface you already use — the same capabilities, governed by conformance (one product, two engines), not a thin re-implementation per surface.
evaluate · Fellegi-Sunter scoring · GoldenFlow transforms. Resolve without moving data out of the warehouse.node:*-free (browsers, Cloudflare Workers, Vercel Edge, Deno); an opt-in WebAssembly backend (await enableWasm()) swaps in the same pyo3-free Rust kernels the Python wheels and SQL UDFs use, with pure-TS as the byte-identical default.Surface parity is not the same as handing any pipeline phase from one language to the other byte-for-byte. Each verdict below is measured by a conformance harness, not assumed:
| Boundary | Verdict |
|---|---|
| Identity graph DB | ✅ byte-safe + cryptographically cross-verifiable (a seal written by one toolkit validates under the other) |
score → cluster and the end-to-end split-run | ✅ byte-safe — reproduces the single-language run |
Cluster JSON · config YAML · Learning Memory · record_fingerprint | ✅ portable |
| String scoring | 🟡 4-decimal tolerance — a pair on a threshold can flip (byte-identical only with the shared WASM scorer) |
| Standardize / dates · embeddings · auto-config controller | 🟠 divergent — not byte-portable |
| Distributed / Ray · document (VLM) ingest | ⛔ Python-only by architecture |
Rule of thumb: hand off at the cluster or identity boundary and it's seamless; don't split across standardize/dates, embeddings, or the controller and expect bit-exact reproduction. Full detail + the runnable harness that keeps these verdicts honest: Cross-language parity & phase-handoff limits.
GoldenMatch is the headline, but resolution is only as good as what feeds it. Five sibling tools clean, standardize, and map records before they reach the identity layer — each stands alone, but they compose into one pipeline, orchestrated declaratively by GoldenPipe:
flowchart LR
raw([raw rows])
golden([golden entities])
subgraph orchestration ["GoldenPipe orchestrates"]
direction LR
infermap[InferMap] --> goldencheck[GoldenCheck] --> goldenflow[GoldenFlow] --> goldenmatch[GoldenMatch]
end
raw --> infermap
goldenmatch --> golden
| Package | Lang | Role in the pipeline | Install |
|---|---|---|---|
| InferMap | Python · TS | Schema mapping — auto-aligns columns across heterogeneous sources | pip install infermap · npm i infermap |
| GoldenCheck | Python · TS | Data-quality scanning — encoding, format validation, anomaly detection | pip install goldencheck · npm i goldencheck |
| GoldenFlow | Python · TS | Transforms & standardizers — phone, date, address, categorical | pip install goldenflow · npm i goldenflow |
| GoldenMatch | Python · TS | Zero-config entity resolution → the identity spine. Headline package. | pip install goldenmatch · npm i goldenmatch |
| GoldenAnalysis | Python · TS | Analysis & reporting — any stage's artifacts → a unified AnalysisReport + cross-run regression detection | pip install goldenanalysis · npm i goldenanalysis |
| GoldenPipe | Python · TS | Orchestrator — declarative YAML wiring the steps | pip install goldenpipe · npm i goldenpipe |
| golden-suite | Python | One-line meta-install: the whole suite + native acceleration | pip install golden-suite |
The deepest docs live in packages/python/goldenmatch/README.md (~1,300 lines: full feature list, CLI, architecture, benchmarks).
The suite owns its string-matching primitives instead of renting them — byte-identical drop-in replacements, published on their own so they're usable outside the suite too.
| Library | Replaces | What it is | Install |
|---|---|---|---|
| goldenfuzz | rapidfuzz | Fuzzy-string scorers + the full fuzz.* composite family + one-vs-many extract/cdist. Byte-identical (oracle-fuzzed), faster on short strings. | pip install goldenfuzz · cargo add goldenfuzz-core |
| goldenphonetic | jellyfish | Phonetic encoders — soundex / metaphone / nysiis / match-rating. Byte-identical (6,000-input + 2,500-pair fuzz corpus), pure-Rust zero-dep. | pip install goldenphonetic · cargo add goldenphonetic-core |
| goldenmatch-hnsw | FAISS IndexHNSWFlat | Pure-Rust HNSW approximate-nearest-neighbor index (zero C deps) — powers embedding-based blocking across Python, Rust, and TS/WASM. | pip install goldenmatch-hnsw |
Entity resolution is the stage most GraphRAG pipelines do worst — duplicate surface forms of one entity scatter across documents. Two packages put GoldenMatch's resolution there:
| Package | What it does | Status |
|---|---|---|
| goldenmatch-kg | Drop-in GoldenMatch resolution as the ER stage of existing KG frameworks (neo4j-graphrag, LlamaIndex, Graphiti). | in-repo · not published (by design) |
| goldengraph | Build-your-own-KG from text — text → LLM extraction → GoldenMatch resolution → durable bi-temporal store. Rust engine; ER is the differentiator. | in-repo · first PyPI release pending |
Measured, not asserted (ER-KG-Bench): resolution scores F1 0.602 on the labelled set, ahead of Neo4j-KGBuilder (0.456), neo4j-graphrag (0.403), and MS-GraphRAG / LightRAG / Cognee / mem0 (0.066). A resolved graph also does two things passage-window RAG structurally can't — exact aggregation (size-invariant where RAG recall collapses 0.99 → 0.64 across cluster-size buckets) and temporal as-of (1.000 vs 0.002 on past-date queries).
Every headline number maps back to a single committed runner (scripts/run_benchmarks.py); see docs/reproducing-benchmarks.md for per-number commands, dataset URLs, and expected output with tolerance.
docs/scale-envelope.md) — per-backend ranges (in-memory/bucket to a few M · DuckDB out-of-core to ~50M · Ray distributed ≥ 50M), block-size failure modes, and a decision tree for picking a backend.Verified at the top end: a full 100M-row dedupe on a 5-node Ray cluster in 9.2 min (554 s), 20,000,000 golden records recovered exactly, driver peak 0.36 GB RSS. The default distributed path is recall-complete — duplicates merge correctly no matter how the input is partitioned (blocking-key shuffle scoring + distributed randomized-contraction WCC), and it stays driver-collect-free end to end. Recipe: configs/distributed-100m.yaml.
Three reproducible real-world pipelines run this on public data at scale:
Dedupe a CSV in 30 seconds — zero config, writes <timestamp>_golden.csv:
pip install goldenmatch && goldenmatch dedupe customers.csv
import goldenmatch as gm
result = gm.dedupe("customers.csv") # zero-config
print(result) # DedupeResult(records=5000, clusters=847, match_rate=12.0%)
result.golden.write_csv("deduped.csv")
result = gm.dedupe("customers.csv", # or be explicit
exact=["email"], fuzzy={"name": 0.85, "zip": 0.95}, blocking=["zip"], threshold=0.85)
import { dedupe } from "goldenmatch"; // edge-safe: browsers, Vercel Edge, Workers, Deno
const result = dedupe(rows, { fuzzy: { name: 0.85 }, blocking: ["zip"], threshold: 0.85 });
The whole suite, configured for speed — golden-suite pulls in every package plus the native (Rust) kernels, pinned and defaulted to the perf-optimized config. Native wheels are hard dependencies on purpose: a platform without a wheel fails loudly rather than silently running the slow pure-Python path.
pip install golden-suite
golden-suite doctor # verify every package + native kernel is importable and healthy
golden-suite optimize # repair / re-enable the perf-optimized config
pip install golden-suite[mcp] # + aggregator MCP server (every tool, one endpoint)
pip install golden-suite[all] # everything
Just GoldenMatch — fat optional extras, pay only for what you use (native acceleration is default on common platforms):
pip install goldenmatch # core (CSV in, CSV out) + native
pip install goldenmatch[documents] # + PDF/image ingest (resolve unstructured input)
pip install goldenmatch[embeddings] # + sentence-transformers, FAISS
pip install goldenmatch[llm] # + Claude / OpenAI for LLM boost
pip install goldenmatch[ray] # + Ray distributed backend (50M+ rows)
pip install goldenmatch[postgres] # + Postgres sync (also: [snowflake] [bigquery] [databricks] [salesforce])
pip install goldenmatch[mcp] # + MCP server (also: [agent] A2A, [web] browser workbench)
Web workbench — pip install 'goldenmatch[web]' then goldenmatch serve-ui my-project (opens http://localhost:5050): edit rules with live validation, preview against a sampled slice, label pairs (mirrored into Learning Memory), compare runs.
More: examples/ — Python (quickstart, full pipeline, customer 360, PPRL, MCP client) · TypeScript (quickstart, Vercel Edge, MCP client) · Airflow.
Remote MCP (nothing to install) — hosted on Smithery; connect any MCP client:
{ "mcpServers": { "goldenmatch": { "url": "https://goldenmatch-mcp-production.up.railway.app/mcp/" } } }
Containers — every package ships as a multi-arch image (linux/amd64 + arm64) on GHCR, pull anonymously:
docker run -p 8300:8300 ghcr.io/benseverndev-oss/goldensuite-mcp:latest # one container, every tool
docker run -p 8200:8200 ghcr.io/benseverndev-oss/goldenmatch-mcp:latest # per-package (also goldencheck/goldenflow/goldenpipe/infermap)
docker run -e POSTGRES_PASSWORD=secret ghcr.io/benseverndev-oss/goldenmatch-extensions:latest # Postgres + extension
Airflow — 13 drop-in DAGs at examples/airflow/ (TaskFlow API, Airflow 2.7+ / 3.x; idempotent, marker-protected), grouped by lifecycle stage:
| Group | DAGs |
|---|---|
| Core pipeline | daily_dedupe, incremental_match, warehouse_native (Snowflake), customer_360, identity_graph |
| Privacy | pprl_linkage (two-party PPRL) |
| Onboarding & monitoring | schema_align_and_load, schema_drift_alarm, quality_gate |
| Feedback loop | review_worker, active_learning |
| Operationalize | reverse_etl (Salesforce/HubSpot), backfill |
goldenmatch/
├── packages/
│ ├── python/ goldenmatch · goldencheck · goldenflow · goldenpipe · infermap · goldenanalysis
│ │ goldensuite-mcp (aggregator) · golden-suite (meta) · goldengraph · goldenmatch-kg
│ ├── typescript/ full TS ports (edge-safe cores + WASM) · goldencheck-types
│ ├── rust/extensions/ Postgres pgrx + DuckDB UDFs + native kernels + owned libraries (own Cargo workspace)
│ ├── dbt/goldensuite/ dbt materializations, tests, macros
│ └── actions/goldencheck/ GitHub Action
├── examples/ python · typescript · airflow (drop-in DAGs)
├── context-network/ architecture decisions + design docs (ADRs, the two-engine frame, Customer 360)
├── docs/superpowers/ design specs and implementation plans
└── justfile · pyproject.toml (uv workspace) · pnpm-workspace.yaml (Turborepo) · .github/workflows/ci.yml
packages/rust/extensions/ is itself a Cargo workspace (the postgres crate is excluded for pgrx); Cargo commands run from inside it.packages/typescript/* form a single pnpm + Turborepo workspace.just install # uv sync + per-package npm install + cargo fetch
just test # all languages · just lint · just build
feature/<name> branches; merge via squash PR. Titles: feat: / fix: / docs:.packages/typescript/goldenmatch/tests/parity/ enforces 4-decimal Python ↔ TypeScript scorer parity.context-network/decisions/ and docs/superpowers/specs/.corepack enable # one-time, picks up pnpm@9.15.0
pnpm install
pnpm turbo run build test typecheck # full pipeline (cached after first run)
Windows: enable Developer Mode so pnpm install can create symlinks; if corepack enable needs admin, npm i -g pnpm@9.15.0 is equivalent.
This repo was formed on 2026-05-01 by folding 8 sibling repos into goldenmatch via git filter-repo (full history preserved). Built by Ben Severn. MIT — see LICENSE.