Dataset Explorer

io.github.khanarmaghanrasheed-18v0.2.1Updated Oct 8, 2026

Help your AI explore local data using summaries, Pearson correlation, eta squared, and Cramer's V.

VerifiedSTDIODesktop onlyFiles & StorageData & Analytics

Overview

AI-generated overview

Lets an assistant read local CSV, Excel, JSON, or Parquet files and compute summaries, missing values, duplicates, outliers, and column associations.

What it does
Dataset Explorer runs locally and loads a dataset file you point it at, then returns computed statistics rather than guesses. Tools include get_dataset_overview, dataset_shape, dataset_statistical_summary, inspect_Column, analyze_target, duplicate_finder, analyze_missing_values, find_correlations, detect_outliers, and screen_target_relationships. It picks Pearson correlation, eta squared, or Cramer's V based on the column types being compared. It also exposes a dataset://guide resource and an explore_dataset prompt.
When to use it
Use it when you want an assistant to inspect a data file on your own machine: checking column types and missing values, finding duplicate rows, spotting outliers, or seeing which features are associated with a target column. It is aimed at exploratory analysis, not modeling or causal claims.
Requirements
A local Python 3.10+ runtime; install from PyPI as dataset-explorer-mcp, or run it with uvx without a separate install. An MCP client that supports local stdio servers, such as Claude Desktop, VS Code with Copilot, or Cursor. Dataset files must be readable on the same computer, and large files need enough RAM since each tool loads the dataset into memory.
Before you install
The server reads the dataset files you point it at and leaves the originals unchanged, but it does not restrict which paths can be read. Results may be sent onward to the assistant's AI provider according to that client's settings, so avoid pointing it at sensitive files unless that is acceptable. Statistics show associations only and do not prove causation.

Installation

In SourceWeft

  1. Open Dataset Explorer in the dashboard and add it to a workspace.
  2. Enable the server for the chats that should use its tools.

Desktop only via STDIO. STDIO servers start a local process, so they need the SourceWeft desktop host.

Other MCP clients

Follow the launch instructions in the repository.

README

Dataset Explorer MCP

Dataset Explorer helps your AI assistant understand data files saved on your computer. Ask it to summarize a dataset, check missing values, find repeated rows, or spot unusual patterns. It calculates answers from your file and leaves the original unchanged.

It works with CSV, TSV, Excel, JSON, and Parquet files. Your assistant starts this small local server when needed. You don't need a web app, a hosting account, or a Gemini API key. Your assistant may send the results to its AI provider according to that client's settings.

The connection uses MCP over stdio, which simply means your assistant talks directly to the server running on your computer.

How it finds relationships

The server calculates statistics from your dataset rather than guessing them:

  • Pearson correlation checks how two numeric columns move together.
  • Eta squared compares numeric values across groups, such as scores across categories.
  • Cramer's V measures the relationship between two category columns.

It chooses the method based on the types of columns being compared. These statistics show associations; they don't prove that one feature causes another.

Install

Available on PyPI.

You need Python 3.10 or newer.

sh
python -m pip install dataset-explorer-mcp

The command to start the server is:

sh
dataset-explorer-mcp

Your MCP client normally runs this command for you. If you run it in a terminal, it waits quietly for messages from a client. That is expected.

If you use uv, you can run it without a separate installation:

sh
uvx dataset-explorer-mcp

Connect your assistant

Use an MCP client that supports local stdio servers, such as Claude Desktop, VS Code with Copilot, or Cursor. The settings file differs by client.

For Claude Desktop or Cursor, add this to your MCP configuration:

json
{  "mcpServers": {    "dataset-explorer": {      "command": "uvx",      "args": ["dataset-explorer-mcp"]    }  }}

For VS Code, use .vscode/mcp.json:

json
{  "servers": {    "dataset-explorer": {      "type": "stdio",      "command": "uvx",      "args": ["dataset-explorer-mcp"]    }  }}

If you installed with pip, use "command": "dataset-explorer-mcp" and "args": [] instead. An absolute path to the executable also works. Reload your client after changing its configuration.

Use a local file

Give your assistant the full path to your dataset. For example:

Explore C:/Users/YourName/Downloads/customers.csv. Check missing values and duplicates.

Summarize /home/yourname/data/sales.xlsx and inspect the revenue column.

In /Users/yourname/data/results.parquet, which features are associated with the target column score?

All tools take a path. A direct tool call looks like:

json
{"path": "C:/Users/YourName/Downloads/customers.csv"}

Use forward slashes in Windows paths, or double backslashes when writing JSON. The file must be available on the computer where the server runs.

Supported files and tools

Supported files: CSV, TSV, Excel (.xlsx, .xls), JSON, and Parquet. Excel reads the first worksheet. JSON must contain tabular data that Pandas can read.

ToolWhat it does
get_dataset_overviewLists columns, types, and missing-value counts
dataset_shapeCounts rows and columns
dataset_statistical_summaryCalculates numeric means and medians
inspect_ColumnSummarizes one column; also takes col_name
analyze_targetInspects a target column; also takes target_name
duplicate_finderFinds repeated rows
analyze_missing_valuesReports missing data
find_correlationsFinds related numeric columns; optional threshold defaults to 0.8
detect_outliersFinds unusual numeric values
screen_target_relationshipsCompares features with a target; also takes target

The server also offers the dataset://guide resource and an explore_dataset prompt. These results help you explore data; they don't prove causes or train a model. Large files need enough RAM because each tool loads the dataset into memory.

Troubleshooting

  • Command not found: use the full path to dataset-explorer-mcp, or install uv and use the uvx configuration above.
  • File not found: use an absolute path and check that the server can read it.
  • No tools appear: check your client's server logs and reload its MCP settings.
  • Server seems idle: it is waiting for the MCP client; connect it through your assistant rather than typing questions into the server terminal.
  • Unsupported file: save the data in one of the formats listed above.
  • Missing values in statistics: empty or constant columns may have undefined statistics. Check the overview and missing-value tools first.

Normal server output is reserved for MCP messages. Diagnostics go to stderr, which your client's server logs usually display.

Run from source

sh
git clone https://github.com/khanarmaghanrasheed-18/MCP-Dataset-Explorer.gitcd MCP-Dataset-Explorerpython -m pip install -e ".[dev]"python -m pytest -qpython mcp_server.py

Build the downloadable package with python -m build.

License

MIT. See LICENSE.

Source: README.md at commit c5b6fbd

Tools

0
Tool metadata has not been indexed yet.

Version history

1
  1. v0.2.1LatestOct 8, 2026