Gujarati Lexicon Mcp

io.github.aarshbharatv0.2.1更新於 Oct 8, 2026

Grounded Gujarati dictionary for AI assistants: meanings, synonyms, idioms, inflections

已驗證STDIO僅桌面Knowledge & Memory

概覽

AI 產生的概覽

讓助理在真實詞典資料中查詢古吉拉特語詞彙的含義、同義詞、慣用語和屈折形式。

功能
提供有依據的古吉拉特語詞典,讓助理查詞而不是猜測。工具包括 define(word)(含義、詞性、例句和轉寫)、synonyms(word)(去重後的同義詞)和 idioms(word)(包含該詞的慣用語)。也提供每日一詞資源和 explain_passage 提示,可逐詞解釋古吉拉特語段落。像 ઘરમાં 這樣的屈折形式會還原到詞根,回應中帶有來源標註。
適用情境
適合處理古吉拉特語文本、需要可靠含義、同義詞或慣用語而非模型猜測的情境。也適合逐詞解釋古吉拉特語段落,或確認某個詞是否確實在詞典中。
執行需求
以本機 stdio 程序執行,從 PyPI 套件 gujarati-lexicon-mcp 安裝。需要 uv。需透過環境變數 YOUR_API_KEY 提供 API 金鑰。僅限桌面端,不提供網頁可執行版本。
安裝前請注意
伺服器要求透過環境變數 YOUR_API_KEY 提供金鑰。詞典資料來自 Wiktionary(CC BY-SA)和人工校對集;人工校對集較小,動詞形式覆蓋不全,部分查詢可能不完整。

安裝

在 SourceWeft 中

  1. 開啟 儀表板中的 Gujarati Lexicon Mcp,將其新增到工作區。
  2. 為需要使用其工具的對話啟用該服務。

Desktop only,透過 STDIO。 STDIO 服務會啟動本機處理程序,因此需要 SourceWeft 桌面主機。

其他 MCP 客戶端

參照 儲存庫 中的啟動說明。

README

Gujarati Lexicon MCP Server [PyPI]

A grounded Gujarati dictionary for AI assistants. Connect it to Claude (or any MCP client) and the assistant looks Gujarati words up in real dictionary data instead of guessing: meanings, synonyms, idioms (રૂઢિપ્રયોગ), and inflected forms like ઘરમાં or આંખોમાં.

Why this exists: LLMs are unreliable with Gujarati. They invent meanings, mix up idioms, and produce convincing but wrong usage. This server gives the model facts to stand on and tells it plainly when a word is not in the dictionary.

[Demo]

Example

User: અમે નો અર્થ શું છે?

Claude (after calling define): અમે means "we", the plural of હું (I)…

User: ઘરમાં નો અર્થ?

Claude: ઘરમાં = ઘર + માં → "in the house". The server matched the inflected form to its base word ઘર.

Features

PrimitiveNameWhat it does
Tooldefine(word)Meanings, part of speech, examples, transliteration
Toolsynonyms(word)Synonyms (સમાનાર્થી શબ્દો) from all sources, deduplicated
Toolidioms(word)Idioms (રૂઢિપ્રયોગ) containing the word; never invents any
Resourcelexicon://word-of-the-dayA daily word from the hand-checked set
Promptexplain_passageExplains a Gujarati passage word by word using the tools
  • About 6,800 words with English meanings, from Wiktionary
  • Hand-checked entries with Gujarati meanings and idioms (curated set, growing)
  • Inflection handling: ઘરમાં, ઘરે, આંખોમાં, and અમે resolve to their base words
  • "Did you mean?" suggestions for typos (પાણિ → પાણી)
  • Source attribution in every response, so the assistant can say where a meaning came from

Quick start

Requires uv.

Add this to your Claude Desktop config (Settings → Developer → Edit Config):

json
      "args": ["gujarati-lexicon-mcp"]

Restart Claude Desktop completely (quit from the system tray), open a new chat, and ask about any Gujarati word.

On Windows, if Claude Desktop can't find uvx, use its full path (find it with where uvx), e.g. C:\\Users\\<you>\\.local\\bin\\uvx.exe.

How lookup works

A word goes through three layers, from most to least reliable:

mermaid
flowchart LR    A[Input word] --> B{Exact headword?}    B -- yes --> C{Pure inflected form?}    C -- yes --> D[Follow Wiktionary link<br/>ઘરે → ઘર]    C -- no --> E[Return entry]    D --> E    B -- no --> F[Strip suffixes<br/>માં, નો, ની, ે, ો ...]    F -- match --> E    F -- no match --> G[Not found + suggestions<br/>tell the model not to guess]
  1. Exact match. A real dictionary entry always wins over a guessed base form.
  2. Wiktionary inflection links. Wiktionary records that ઘરે is the locative of ઘર. The converter reads the structured form_of field, and falls back to parsing glosses like "plural of X" where that field is missing.
  3. Suffix stripping. As a last resort, common endings are removed (up to two, e.g. આંખોમાં → આંખો → આંખ). A stripped form is accepted only if it is a real headword, so words are never mangled.

Responses keep sources separate (curated vs wiktionary_senses), so the model can attribute meanings honestly.

Data sources

SourceProvidesLicense
curated.jsonGujarati meanings, idioms, synonyms (hand-checked)MIT (this project)
wiktionary.json~6,800 words: English meanings, transliterations, synonyms, inflection linksCC BY-SA

Wiktionary data comes from Wiktionary via kaikki.org, extracted with wiktextract:

Tatu Ylonen. Wiktextract: Wiktionary as Machine-Readable Structured Data. Proceedings of the 13th Conference on Language Resources and Evaluation (LREC), pp. 1317–1325, 2022.

Rebuilding the Wiktionary data

The raw dump is not committed. To regenerate wiktionary.json:

bash
mkdir -p data/rawcurl -L -o data/raw/kaikki-gujarati.jsonl "https://kaikki.org/dictionary/Gujarati/kaikki.org-dictionary-Gujarati.jsonl"uv run python scripts/build_wiktionary.py

The converter groups entries by word, keeps only what the tools need (20 MB → a compact JSON file), drops non-Gujarati-script headwords, and records inflection links.

Development

bash
git clone https://github.com/aarshbharat/gujarati-lexicon-mcpcd gujarati-lexicon-mcpuv syncuv run pytest -v                                    # run testsuv run mcp dev src/gujarati_lexicon_mcp/server.py   # open MCP Inspector

Project layout:

src/gujarati_lexicon_mcp/├── server.py           # MCP server: tools, resource, prompt└── data/               # dictionary data shipped with the packagescripts/build_wiktionary.py   # raw kaikki dump → wiktionary.jsontests/                        # pytest

Known limitations

  • Verb forms are only partly covered: forms Wiktionary links (e.g. હોઈશ → હોવું) work, but phrases like લે છે are not lemmatized.
  • Wiktionary synonyms are merged across senses. ઘર's list includes ઓફિસ (from its "office" sense). The tool tells the model this.
  • The curated set is small. Most entries have English meanings only; Gujarati-language meanings and idioms come from the hand-checked data.
  • Grounding covers facts, not everything the model says. The model may add correct background from its own knowledge (e.g. the inclusive/exclusive "we" distinction for અમે/આપણે). The server's instructions ask it to label general knowledge, but cannot force it.

[PyPI]

Roadmap

  • Grow the curated set, especially idioms and proverbs (કહેવત)
  • Sense-level synonyms instead of a merged list
  • Better verb lemmatization
  • Streamable HTTP transport and a hosted endpoint
  • Publish to PyPI and the MCP Registry

License

  • Code: MIT
  • Wiktionary-derived data (src/gujarati_lexicon_mcp/data/wiktionary.json): CC BY-SA, see Wiktionary:Copyrights

Built by Aarsh Dhokai · AI × AI: Artificial Intelligence × Aarsh India

來源:README.md,提交 a92d798

工具

0
工具後設資料尚未被收錄。

版本歷史

1
  1. v0.2.1最新Oct 8, 2026