
Dataset Explorer
io.github.khanarmaghanrasheed-18v0.2.1Updated Oct 8, 2026
Help your AI explore local data using summaries, Pearson correlation, eta squared, and Cramer's V.
Overview
Lets an assistant read local CSV, Excel, JSON, or Parquet files and compute summaries, missing values, duplicates, outliers, and column associations.
- What it does
- Dataset Explorer runs locally and loads a dataset file you point it at, then returns computed statistics rather than guesses. Tools include get_dataset_overview, dataset_shape, dataset_statistical_summary, inspect_Column, analyze_target, duplicate_finder, analyze_missing_values, find_correlations, detect_outliers, and screen_target_relationships. It picks Pearson correlation, eta squared, or Cramer's V based on the column types being compared. It also exposes a dataset://guide resource and an explore_dataset prompt.
- When to use it
- Use it when you want an assistant to inspect a data file on your own machine: checking column types and missing values, finding duplicate rows, spotting outliers, or seeing which features are associated with a target column. It is aimed at exploratory analysis, not modeling or causal claims.
- Requirements
- A local Python 3.10+ runtime; install from PyPI as dataset-explorer-mcp, or run it with uvx without a separate install. An MCP client that supports local stdio servers, such as Claude Desktop, VS Code with Copilot, or Cursor. Dataset files must be readable on the same computer, and large files need enough RAM since each tool loads the dataset into memory.
Installation
In SourceWeft
- Open Dataset Explorer in the dashboard and add it to a workspace.
- Enable the server for the chats that should use its tools.
Desktop only via STDIO. STDIO servers start a local process, so they need the SourceWeft desktop host.
Other MCP clients
Follow the launch instructions in the repository.
README
Dataset Explorer MCP
Dataset Explorer helps your AI assistant understand data files saved on your computer. Ask it to summarize a dataset, check missing values, find repeated rows, or spot unusual patterns. It calculates answers from your file and leaves the original unchanged.
It works with CSV, TSV, Excel, JSON, and Parquet files. Your assistant starts this small local server when needed. You don't need a web app, a hosting account, or a Gemini API key. Your assistant may send the results to its AI provider according to that client's settings.
The connection uses MCP over stdio, which simply means your assistant talks directly to the server running on your computer.
How it finds relationships
The server calculates statistics from your dataset rather than guessing them:
- Pearson correlation checks how two numeric columns move together.
- Eta squared compares numeric values across groups, such as scores across categories.
- Cramer's V measures the relationship between two category columns.
It chooses the method based on the types of columns being compared. These statistics show associations; they don't prove that one feature causes another.
Install
Available on PyPI.
You need Python 3.10 or newer.
The command to start the server is:
Your MCP client normally runs this command for you. If you run it in a terminal, it waits quietly for messages from a client. That is expected.
If you use uv, you can run it without a separate installation:
Connect your assistant
Use an MCP client that supports local stdio servers, such as Claude Desktop, VS Code with Copilot, or Cursor. The settings file differs by client.
For Claude Desktop or Cursor, add this to your MCP configuration:
For VS Code, use .vscode/mcp.json:
If you installed with pip, use "command": "dataset-explorer-mcp" and "args": []
instead. An absolute path to the executable also works. Reload your client after
changing its configuration.
Use a local file
Give your assistant the full path to your dataset. For example:
Explore
C:/Users/YourName/Downloads/customers.csv. Check missing values and duplicates.
Summarize
/home/yourname/data/sales.xlsxand inspect the revenue column.
In
/Users/yourname/data/results.parquet, which features are associated with the target columnscore?
All tools take a path. A direct tool call looks like:
Use forward slashes in Windows paths, or double backslashes when writing JSON. The file must be available on the computer where the server runs.
Supported files and tools
Supported files: CSV, TSV, Excel (.xlsx, .xls), JSON, and Parquet.
Excel reads the first worksheet. JSON must contain tabular data that Pandas can read.
The server also offers the dataset://guide resource and an explore_dataset
prompt. These results help you explore data; they don't prove causes or train a model.
Large files need enough RAM because each tool loads the dataset into memory.
Troubleshooting
- Command not found: use the full path to
dataset-explorer-mcp, or install uv and use theuvxconfiguration above. - File not found: use an absolute path and check that the server can read it.
- No tools appear: check your client's server logs and reload its MCP settings.
- Server seems idle: it is waiting for the MCP client; connect it through your assistant rather than typing questions into the server terminal.
- Unsupported file: save the data in one of the formats listed above.
- Missing values in statistics: empty or constant columns may have undefined statistics. Check the overview and missing-value tools first.
Normal server output is reserved for MCP messages. Diagnostics go to stderr, which your client's server logs usually display.
Run from source
Build the downloadable package with python -m build.
License
MIT. See LICENSE.
Source: README.md at commit c5b6fbd
Tools
0Version history
1- v0.2.1LatestOct 8, 2026


