jDataMunch is an MCP server for coding agents and analysts that answers questions about CSV, Excel, Parquet, and JSONL files without pasting the rows into the context window.
Index a dataset once, then retrieve column profiles, filtered rows, server-side aggregations, and cross-dataset joins — so a million-row file costs thousands of tokens instead of millions.
Install · Quickstart · Benchmarks · Commercial licensing
Free for personal use. Commercial use requires a paid license — terms below.
The problem. The default way an agent explores a spreadsheet is to paste it into the prompt. A 255 MB CSV with a million rows costs roughly 111 million tokens that way, and the model still has to reason through a million rows to answer "what columns are in here?"
The mechanism. jDataMunch profiles the file once — columns, types, cardinality, null rates, distributions — and stores that locally. Queries then run against the data, not against a copy of it in the prompt: filters, aggregations, and joins execute server-side and return only results.
The outcome. Orientation questions are answered from the profile. Row-level questions return matching rows. The raw file never enters the context window.
Measured on a real public dataset, not estimated. Full harness and per-query results in benchmarks/.
Corpus: LAPD crime records — 1,004,894 rows, 28 columns, 255 MB Baseline: 111,028,360 tokens to paste the raw file
describe_dataset: ~3,849 tokens — a 25,333× reduction Methodology & harness · Full results
| Task | Without jDataMunch | With jDataMunch | Reduction |
|---|---|---|---|
| Understand a dataset's shape | Paste 111M tokens | describe_dataset → ~3,849 tokens |
~25,000× |
| Schema + one column deep-dive | Paste 111M tokens | describe_dataset + describe_column → ~4,400 tokens |
~25,000× |
| Filter to matching rows | Load all 1M rows | get_rows with filters → matching rows only |
~99%+ |
| Count by category | Return all rows, aggregate in the model | aggregate(group_by=[...]) → 21 rows |
~99.9% |
What these numbers are and are not. The reduction is measured against pasting the complete file, which is what a naive agent does and what the token bill reflects. It is not measured against a competent human analyst who would never paste a 255 MB CSV. The multiple scales with file size: a 200-row spreadsheet has far less to save, and the honest figure there is closer to "no meaningful difference."
Typical latencies from the same run: describe_column on a single column, 22–33 ms and ~600 tokens.
Requirements: Python 3.10+, any MCP-compatible client.
There is no install step. jdatamunch-mcp is a stdio MCP server with no CLI subcommands, so nothing needs to land on your PATH — point your client at uvx and it fetches and runs the server on demand.
Claude Code setup:
claude mcp add jdatamunch -- uvx jdatamunch-mcpNothing else. Don't have uv yet?
Reading Excel or Parquet? Those pull optional extras, which uvx takes on the --from argument:
claude mcp add jdatamunch -- uvx --from "jdatamunch-mcp[excel,parquet]" jdatamunch-mcpPrefer a persistent install?
| Command | Use it when |
|---|---|
uv tool install jdatamunch-mcp |
You want it resolved once instead of per-launch |
pipx install jdatamunch-mcp |
You already standardise on pipx |
pip install jdatamunch-mcp |
Inside a virtualenv you manage yourself. ⚠ Refused on PEP 668 distros (Ubuntu 24.04+, Debian 12+) — use one of the two above. |
Extras take the usual bracket form here: uv tool install "jdatamunch-mcp[excel,parquet]". Registering the server still works the same way; substitute jdatamunch-mcp for uvx jdatamunch-mcp in the claude mcp add line above.
Restart Claude Code, then type /mcp — jdatamunch should be listed. That listing is the verification step; running the server directly just waits on stdin.
Full per-client setup, including Claude Desktop, Cursor, and Windsurf: QUICKSTART.md.
Assumes: jDataMunch installed and registered with your client, and a CSV to hand.
Everything happens inside your agent — there is no separate indexing command. Ask it to index:
Using jdatamunch, index ./data/sales.csv
It calls index_local, which returns the dataset name, row and column counts, and detected types. Then:
Using jdatamunch, describe the sales dataset and tell me which columns have missing values.
The agent calls describe_dataset, which returns column names, inferred types, cardinality, null rates, and sample values — without reading a single row into context. _meta.tokens_saved reports what that cost against loading the file.
Next step: describe_column for a distribution on one column, or aggregate to group and count server-side.
- Orient in a dataset you have never seen.
describe_dataset,describe_column,sample_rows,get_distribution,get_correlations. - Query without loading rows.
get_rowswith filters,aggregatewithgroup_by,run_sql, andplan_queryto preview cost before running. - Work across datasets.
suggest_joins,suggest_keys,join_datasets. - Find data-quality problems.
get_dataset_health,data_health_radar,get_data_hotspots(null rate, cardinality anomalies, outlier spread),get_schema_drift,find_unused_columns. - Preflight schema changes.
check_column_drop_safeandget_schema_impactbefore you drop or rename. - Search semantically.
search_dataandfind_similar_columnswhen you know what you mean but not what it is called. - Index from GitHub.
index_repopulls CSV, Excel, Parquet, and JSONL straight from a repository, incrementally by HEAD SHA, private repos included.
39 tools in total. Full reference: USER-MANUAL.md.
Everything runs locally. The dataset is profiled on your machine and the index is stored on your machine; no hosted service is involved in indexing or querying.
data.csv ──► profiler ──► column stats + local index
│
MCP client ◄── query ─┘ (filters, aggregates, joins
execute server-side)
Aggregations and filters run against the stored data rather than being simulated in the model, which is why the row count barely affects the token cost of an answer. Sampling-based statistics report their error bounds (roughly 2% standard error) rather than presenting an estimate as exact.
| Format | Extensions | Install extra |
|---|---|---|
| CSV / TSV | .csv, .tsv |
built in |
| JSON Lines | .jsonl |
built in |
| Excel | .xlsx, .xls |
pip install "jdatamunch-mcp[excel]" |
| Parquet | .parquet |
pip install "jdatamunch-mcp[parquet]" |
Local-first. Your data is profiled and indexed on your machine and is not uploaded.
The base package's only default network behavior is an anonymous savings counter — a random ID plus aggregate token counts. No data, no column names, no file paths, no PII. Opt out completely:
JDATAMUNCH_SHARE_SAVINGS=0index_repo reaches GitHub only when you invoke it, using a token you supply. Embedding providers are called only when you configure one. There is no scheduler and no background reporting.
Full detail, including what each optional extra pulls in: SECURITY.md.
- Savings scale with file size. On a small spreadsheet the difference is negligible; the benchmark figures come from a 255 MB file.
- Sampled statistics are sampled. Distribution and correlation figures on very large files carry a stated error bound rather than being exact.
- Excel and Parquet need optional extras, which pull additional dependencies.
- A default
describe_columnwill not be labelledoffloadable. jDataMunch does not assert index freshness it cannot prove, so the cheap freshness reading answersunknownand the annotation fails closed. That is deliberate — see the annotation section. - jDataMunch does not read code or prose. Code symbols belong to jcodemunch-mcp; documentation sections to jdocmunch-mcp.
JMUNCH_OFFLOADABLE=1 (suite-wide) or JDATAMUNCH_OFFLOADABLE=1 (this server only) makes describe_column carry an advisory _meta.offloadable block marking whether the answer is simple and self-contained enough to hand to a cheaper model.
It is a label and nothing else. jDataMunch never calls another model, never routes the request, and never touches your API keys. Off by default; you decide what happens next.
The verdict is tri-state and reason-coded: not_evaluated ("we did not assess it") is not not_offloadable ("this is not simple work"). It fails closed — any unknown bearing on the answer disqualifies, because a false offloadable sends real work to a model that will confabulate over the gap. verify_with names the call that would adjudicate a cheaper model's answer.
Identical field contract across all three jMunch servers, with a pinned contract digest that fails the build in any one of them that drifts.
| Doc | What it covers |
|---|---|
| QUICKSTART.md | Zero-to-indexed in three steps |
| USER-MANUAL.md | Full guide for analysts, ops, and non-developers |
| SECURITY.md | Data handling, network behavior, vulnerability reporting |
| benchmarks/METHODOLOGY.md | How the benchmark is run and what it measures |
| CONTRIBUTING.md | Development setup and the CLA requirement |
| CHANGELOG.md | Release history |
Released under the jDataMunch-MCP Dual-Use License (full terms). Free for non-commercial use. Commercial use requires a paid license, one-time, sold by jMunch LLC.
jDataMunch only: Builder, $39 (1 developer) · Studio, $149 (up to 5) · Platform, $499 (org-wide internal deployment)
Full jMunch suite (code + docs + data): Trio Builder, $99 · Trio Studio, $449 · Trio Platform, $2,499
Individual developers and non-commercial projects need no license. Organizations deploying jDataMunch across internal teams do.
Actively maintained. Issues and bug reports: GitHub Issues. Commercial licensing questions go through jcodemunch.com.
Part of the jMunch suite alongside jcodemunch-mcp (code symbols) and jdocmunch-mcp (documentation sections). All three implement jMRI, the open retrieval interface spec — same response envelope, same token accounting.