
Gujarati Lexicon Mcp
io.github.aarshbharatv0.2.1更新于 Oct 8, 2026
Grounded Gujarati dictionary for AI assistants: meanings, synonyms, idioms, inflections
概览
让助手在真实词典数据中查询古吉拉特语单词的含义、同义词、习语和屈折形式。
- 功能
- 提供有依据的古吉拉特语词典,让助手查词而不是猜测。工具包括 define(word)(含义、词性、例句和转写)、synonyms(word)(去重后的同义词)和 idioms(word)(包含该词的习语)。还提供每日一词资源和 explain_passage 提示,可逐词解释古吉拉特语段落。像 ઘરમાં 这样的屈折形式会还原到词根,响应中带有来源标注。
- 适用场景
- 适合处理古吉拉特语文本、需要可靠含义、同义词或习语而非模型猜测的场景。也适合逐词解释古吉拉特语段落,或确认某个词是否确实在词典中。
- 运行要求
- 作为本地 stdio 进程运行,从 PyPI 包 gujarati-lexicon-mcp 安装。需要 uv。需要通过环境变量 YOUR_API_KEY 提供 API 密钥。仅限桌面端,不提供网页可执行版本。
安装
在 SourceWeft 中
- 打开 控制台中的 Gujarati Lexicon Mcp,将其添加到工作区。
- 为需要使用其工具的对话启用该服务。
Desktop only,通过 STDIO。 STDIO 服务会启动本地进程,因此需要 SourceWeft 桌面宿主。
其他 MCP 客户端
参照 仓库 中的启动说明。
README
Gujarati Lexicon MCP Server [PyPI]
A grounded Gujarati dictionary for AI assistants. Connect it to Claude (or any MCP client) and the assistant looks Gujarati words up in real dictionary data instead of guessing: meanings, synonyms, idioms (રૂઢિપ્રયોગ), and inflected forms like ઘરમાં or આંખોમાં.
Why this exists: LLMs are unreliable with Gujarati. They invent meanings, mix up idioms, and produce convincing but wrong usage. This server gives the model facts to stand on and tells it plainly when a word is not in the dictionary.
Example
User: અમે નો અર્થ શું છે?
Claude (after calling
define): અમે means "we", the plural of હું (I)…
User: ઘરમાં નો અર્થ?
Claude: ઘરમાં = ઘર + માં → "in the house". The server matched the inflected form to its base word ઘર.
Features
- About 6,800 words with English meanings, from Wiktionary
- Hand-checked entries with Gujarati meanings and idioms (curated set, growing)
- Inflection handling: ઘરમાં, ઘરે, આંખોમાં, and અમે resolve to their base words
- "Did you mean?" suggestions for typos (પાણિ → પાણી)
- Source attribution in every response, so the assistant can say where a meaning came from
Quick start
Requires uv.
Add this to your Claude Desktop config (Settings → Developer → Edit Config):
Restart Claude Desktop completely (quit from the system tray), open a new chat, and ask about any Gujarati word.
On Windows, if Claude Desktop can't find
uvx, use its full path (find it withwhere uvx), e.g.C:\\Users\\<you>\\.local\\bin\\uvx.exe.
How lookup works
A word goes through three layers, from most to least reliable:
- Exact match. A real dictionary entry always wins over a guessed base form.
- Wiktionary inflection links. Wiktionary records that ઘરે is the locative of ઘર. The converter reads the structured
form_offield, and falls back to parsing glosses like "plural of X" where that field is missing. - Suffix stripping. As a last resort, common endings are removed (up to two, e.g. આંખોમાં → આંખો → આંખ). A stripped form is accepted only if it is a real headword, so words are never mangled.
Responses keep sources separate (curated vs wiktionary_senses), so the model can attribute meanings honestly.
Data sources
Wiktionary data comes from Wiktionary via kaikki.org, extracted with wiktextract:
Tatu Ylonen. Wiktextract: Wiktionary as Machine-Readable Structured Data. Proceedings of the 13th Conference on Language Resources and Evaluation (LREC), pp. 1317–1325, 2022.
Rebuilding the Wiktionary data
The raw dump is not committed. To regenerate wiktionary.json:
The converter groups entries by word, keeps only what the tools need (20 MB → a compact JSON file), drops non-Gujarati-script headwords, and records inflection links.
Development
Project layout:
Known limitations
- Verb forms are only partly covered: forms Wiktionary links (e.g. હોઈશ → હોવું) work, but phrases like લે છે are not lemmatized.
- Wiktionary synonyms are merged across senses. ઘર's list includes ઓફિસ (from its "office" sense). The tool tells the model this.
- The curated set is small. Most entries have English meanings only; Gujarati-language meanings and idioms come from the hand-checked data.
- Grounding covers facts, not everything the model says. The model may add correct background from its own knowledge (e.g. the inclusive/exclusive "we" distinction for અમે/આપણે). The server's instructions ask it to label general knowledge, but cannot force it.
Roadmap
- Grow the curated set, especially idioms and proverbs (કહેવત)
- Sense-level synonyms instead of a merged list
- Better verb lemmatization
- Streamable HTTP transport and a hosted endpoint
- Publish to PyPI and the MCP Registry
License
- Code: MIT
- Wiktionary-derived data (
src/gujarati_lexicon_mcp/data/wiktionary.json): CC BY-SA, see Wiktionary:Copyrights
Built by Aarsh Dhokai · AI × AI: Artificial Intelligence × Aarsh India
来源:README.md,提交 a92d798
工具
0版本历史
1- v0.2.1最新Oct 8, 2026


