
DesiData
io.github.krishnakaushik195v0.1.0Updated Oct 9, 2026
Search India datasets; inspect sources, licences and previews; get notebook and Python examples.
Overview
Lets an assistant search India-focused datasets, inspect their metadata and licences, preview samples, and get notebook or Python loader examples.
- What it does
- DesiData is a remote MCP service that exposes the published DesiData catalogue of India-focused datasets. An assistant can search the catalogue, inspect dataset metadata and provenance, preview a small sample of a dataset, and retrieve notebook or Python loader examples. Search and previews are read-only and need no account or token.
- When to use it
- Use it when you want an assistant to find Indian datasets, check what a dataset contains and how it is licensed, or get ready-made loader code before downloading. It is aimed at data exploration and notebook workflows rather than full data retrieval.
- Requirements
- A remote streamable HTTP endpoint; no package install is declared. Search and previews need no account or token. Full dataset downloads happen outside the MCP flow and require a free DesiData account and a DD_TOKEN environment variable or Colab secret.
Installation
In SourceWeft
- Open DesiData in the dashboard and add it to a workspace.
- Enable the server for the chats that should use its tools.
Web executable via Streamable HTTP. Remote servers run from the web runtime once configured in a workspace.
Other MCP clients
Add this to your client's mcpServers config.
{
"mcpServers": {
"desidata-mcp": {
"type": "http",
"url": "https://lqcxdxjncdzzkwgveydz.supabase.co/functions/v1/mcp"
}
}
}README
DesiData
DesiData publishes India-focused datasets and tools for exploring and working with them. This repository contains dataset notebooks and examples for using DesiData data in Google Colab and Python.
- Website: desidata.in
- Browse datasets: DesiData datasets
- Weekly model benchmark: DesiData Bench
- Community announcements and discussion: DesiData Community
- Python package: desidata on PyPI
Use a dataset notebook
Open a notebook under notebooks/ and run it in Colab or your own Python environment. The notebook collection is refreshed as datasets are added.
Some notebook download examples may need updating to use the current authenticated download flow. See Download data with a DD token before running a download cell.
Download data with a DD token
DesiData datasets are free. To associate programmatic downloads with your DesiData account, create a free DD token from your DesiData profile.
- Sign in to DesiData and create a token in your profile.
- Copy and save it when it is shown. Treat it like a password; do not paste it into a notebook that you share or commit it to GitHub.
- Set it as the
DD_TOKENenvironment variable in your local environment, or add it to Colab's Secrets asDD_TOKEN. - Use the token in the DesiData Python package or authenticated download examples.
For example, in a local terminal:
In Google Colab, add DD_TOKEN under the notebook's Secrets panel and grant the notebook access to it. Do not put the token directly in a notebook cell.
A token is free. Authenticated download requests are associated with your account so DesiData can show download activity and counts. The count records a download request; it does not prove that a file finished downloading or was used.
Weekly benchmark
The DesiData benchmark page compares selected open models on the current benchmark set. The plan is to run and publish a refreshed model comparison each Saturday afternoon, using the latest DesiData benchmark data. Check DesiData Bench for the current results and methodology.
Updates
This repository is a public place to follow notebook and data-access changes. New dated entries will be added here as the project changes.
2026-09-27 — DD token downloads
Programmatic dataset downloads now use a free DD token linked to a DesiData account. This lets download activity be attributed to the account while keeping the datasets free. Create a token from your profile; never commit it to a notebook or repository.
Planned — Weekly model benchmark
The plan is to publish updated open-model benchmark results on Saturday afternoons. Results and details will be posted on the benchmark page.
Contributing and feedback
Have a dataset request, notebook correction, or benchmark suggestion? Open a GitHub issue or join the conversation on the DesiData Community page.
DesiData MCP plugin
The DesiData MCP plugin lets AI assistants search the published catalogue, inspect dataset metadata and provenance, preview a small sample, and get notebook or Python loader examples. Search and previews are read-only and do not need an account or API token. Full dataset downloads continue to use the user's own DesiData DD_TOKEN through the normal download flow.
The same remote MCP service works with Cursor, Codex, and Gemini CLI. See the plugin setup guide for connection instructions and current limits.
Install in Gemini CLI
Restart Gemini CLI after installation, then run /mcp list to confirm DesiData is connected. Search and previews work without signing in; full dataset downloads use the user's own DesiData DD_TOKEN through the regular download flow.
Source: README.md at commit 3bd7dc6
Tools
0Version history
1- v0.1.0LatestOct 9, 2026

