PersonaMCP

io.github.robyrorov0.1.0更新於 Oct 1, 2026

Local-first memory of how you write: style profile and real reply examples from your chat exports

概覽

AI 產生的概覽

PersonaMCP 在本機匯入你的聊天匯出,量測你的寫作風格,並向 AI 助理提供風格設定檔與真實回覆範例。

功能
它將 Instagram、Snapchat、WhatsApp 或通用 JSON/JSONL 對話匯出匯入本機 SQLite,並只分析你送出的文字,統計長度、大小寫、標點、表情符號、詞彙與語言標記。get_style_profile、get_person_style、search_messages、find_similar_interactions、get_writing_context 與 get_persona_summary 等工具會回傳可量測的風格證據,以及有限數量的歷史來訊/回覆配對。選用的語意檢索完全在本機執行,提供以向量為基礎的相似度。它不產生回覆,由連線的助理負責撰寫。
適用情境
當你希望助理依照你真實的寫作習慣草擬訊息,或檢索你過去對某個人、某種情境如何回覆時使用。適合已有聊天匯出、想要有憑據的風格而非「隨性一點」這類模糊指示的使用者。
執行需求
需要 Python 3.11 或更新版本,從 PyPI 安裝或以 uvx 執行;它是以本機 stdio 程序由 MCP 用戶端啟動。必須先匯入聊天匯出,且本人身分需與匯出中的傳送者名稱或 ID 完全相符。選用的語意檢索會從 Hugging Face 下載模型權重。未宣告需要 API 金鑰或帳號。
安裝前請注意
匯入的聊天與產生的設定檔屬於敏感資料;SQLite 未加密,請使用私人資料目錄並啟用磁碟加密。若助理使用外部 LLM,工具結果可能傳送給該服務商。刪除某個人會移除整段對話,包括包含該人的群組。歷史內容以不受信任的資料形式回傳,不應被當成指令執行。

安裝

在 SourceWeft 中

  1. 開啟 儀表板中的 PersonaMCP,將其新增到工作區。
  2. 為需要使用其工具的對話啟用該服務。

Desktop only,透過 STDIO。 STDIO 服務會啟動本機處理程序,因此需要 SourceWeft 桌面主機。

其他 MCP 客戶端

參照 儲存庫 中的啟動說明。

README

PersonaMCP

Your communication style and memory for AI agents.

PersonaMCP imports your conversation exports, measures how you write, and gives an AI agent a compact style profile plus relevant incoming-message/reply examples. Your raw history stays on your computer. There is no model training, web dashboard, cloud embedding requirement, telemetry, or automatic reply generation.

Writing preferences are usually too vague: “sound casual” does not capture someone who uses lowercase, sends three short messages, switches languages, or writes differently to a colleague. PersonaMCP gives the connected writer evidence instead of a guessed personality.

Install

Python 3.11 or newer. Install from PyPI:

sh
python -m pip install personamcppersona --help

Or run it without installing, using uv: uvx personamcp --help.

To work from a checkout of this repository:

sh
git clone https://github.com/robyroro/PersonaMCP.gitcd PersonaMCPuv sync --lockeduv run persona --help

Or install into your own virtual environment:

sh
python -m venv .venv# Linux/macOS: source .venv/bin/activate# Windows PowerShell: .venv\Scripts\Activate.ps1python -m pip install .persona --help

In a uv checkout, prefix the persona commands below with uv run. With an activated pip installation, run persona directly.

The base installation supports import, analysis, FTS search, lexical interaction retrieval, and MCP. Semantic retrieval is an optional, fully local extra:

After initialization and imports (described below):

sh
uv sync --locked --extra semanticuv run persona model prepareuv run persona index

With pip, use python -m pip install '.[semantic]'. Preparing the model explicitly downloads weights from Hugging Face and pins their immutable revision. It does not read or send chats. Indexing and subsequent queries load only local files with remote code disabled.

Quick start

sh
persona initpersona config set-name "Robert"persona config add-alias "roby"persona config set-name "Exact Instagram display name" --platform instagrampersona config set-name "snapchat_username" --platform snapchat
persona import instagram ./instagram-export/persona import snapchat ./snapchat-export/persona import whatsapp ./chat.txtpersona import json ./messages.json
persona statspersona analyzepersona search "cat costa"persona similar "mai vii azi?" --person "David"persona writing-context "ce faci diseara?" --person "David" --platform instagram

For a synthetic first run, use a separate data directory:

sh
persona --home ./sample-persona initpersona --home ./sample-persona config set-name "Owner"persona --home ./sample-persona import json ./examples/messages.jsonpersona --home ./sample-persona analyzepersona --home ./sample-persona writing-context "mai vii azi la cafea?" --person "Alex"

Put private imports and custom data directories outside your repository. The included .gitignore covers conventional private folders but cannot protect every arbitrarily named path.

Identity must match an exact exported sender name or sender ID. Platform aliases override global names on that platform. No participant is guessed from message volume. An import matching no owner messages fails before writing anything. Changing names/aliases recomputes ownership and interactions and invalidates profiles/vectors; run analyze and index again.

Supported imports

PlatformSupported inputNotes
Instagrammessage_N.json or current Meta message_N.html foldersPagination is merged. Common broken JSON Unicode is repaired. HTML scripts/links/media are never executed.
Snapchatchat_history.json, supported chat-card HTML subpages, or original ZIP exportsDirect recipient-keyed and nested JSON layouts. ZIPs are read without extraction; JSON takes precedence over duplicate HTML.
WhatsAppUTF-8 TXT, Android and bracketed iOS timestampsSlash-separated dates; day/month default, --month-first for US exports. Multiline text is retained; system/media notices are excluded from style.
GenericJSON conversation object, conversations object, message array, or JSONLSee the schema below.

Only chat files are imported. Media is not opened, transcribed, downloaded, or analyzed. Unsupported variants fail clearly. A malformed file rolls back the import as a whole. Original files stay untouched. Content hashes and stable message IDs prevent repeated imports from duplicating rows.

WhatsApp has no stable export thread ID. By default the filename identifies the conversation; use --conversation-id "stable-chat-name" when importing renamed or refreshed exports. Offsetless timestamps use a documented UTC convention for wall-clock ordering; they are not claimed to have been recorded in UTC. Instagram/Snapchat HTML support English export dates.

Generic JSON:

json
{  "id": "stable-conversation-id",  "platform": "json",  "title": "Alex",  "participants": ["Owner", "Alex"],  "context": "casual",  "messages": [    {"id": "external-message-id", "sender": "Alex", "sender_id": "account-123",     "text": "mai vii azi?", "timestamp": "2024-01-01T10:00:00Z"},    {"sender": "Owner", "text": "da gen vin acu", "timestamp": "2024-01-01T10:00:10Z"}  ]}

For JSONL, each line is a message with sender, text, timestamp, and an optional conversation_id. Optional external IDs and reply_to are preserved. Generic identity comes from configured aliases, never an imported is_user assertion. Use explicit conversation IDs when importing different datasets; absent IDs use a documented default conversation.

What is measured

persona analyze writes persona.md and persona-profile.json in the private data directory and stores structured profiles in SQLite. Only outgoing text feeds the analyzer. Incoming text is kept as bounded retrieval context.

Measurements include character/word/sentence lengths, short message bursts, casing, punctuation, emoji codepoints, repeated characters, common words and phrases, repeated slang/abbreviation forms, greetings, sign-offs, Romanian diacritics, and Romanian/English word markers. Language markers are heuristics, not a language classifier. Emoji counts measure codepoints, not complete grapheme clusters. Repeated phrases/forms need at least three observations; isolated misspellings are not instructions to add typos.

Context labels are explicit and extensible:

sh
persona conversationspersona set-context CONVERSATION_ID businesspersona set-context OTHER_ID casualpersona analyzepersona person "David"

Global profiles always use available outgoing text. Separate context/person profiles require at least 20 messages by default. On-demand person statistics report their evidence count. No family, dating, personality, sarcasm, or psychological classification is inferred. Recipient-specific queries conservatively exclude groups. Exact names may occur on multiple platforms; pass a platform to retrieval/writing-context when that distinction matters.

Search and semantic retrieval

SQLite FTS5 searches outgoing messages and incoming interaction contexts. Interactions retain up to three incoming messages and a burst of up to eight outgoing replies. A two-hour gap or a media record breaks pairing; an outgoing burst spans at most five minutes between messages.

The local semantic provider uses sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2. It embeds incoming contexts and stores float32 vectors in SQLite. Queries use exact cosine scans over eligible vectors; this favors a simple local architecture over an external vector server. Indexes resume after interruption. A recipient/context/platform filter applies before ranking.

If eligible vectors are absent, retrieval reports lexical explicitly. A partially built index reports its coverage. Semantic scores below 0.25 are omitted; these scores are not probabilities or guarantees of relevance. There is no silent cross-recipient fallback. Future providers can implement the EmbeddingProvider protocol without changing the import or writing interfaces.

MCP setup

The server uses the official Python MCP SDK and stdio transport. Stdout carries only protocol messages; diagnostic output goes to stderr.

sh
persona serve

If the data directory has not been initialized yet, serve creates it the same way persona init does and reports this on stderr; tools return empty results until you import exports. Usually the client launches this command for you. Use absolute paths because the client's working directory can differ from your terminal's. Example configuration for clients accepting the common mcpServers structure:

json
{  "mcpServers": {    "personamcp": {      "command": "/absolute/path/to/venv/bin/persona",      "args": ["--home", "/absolute/path/to/private/persona-data", "serve"]    }  }}

With uv installed, the client can run the published package directly:

json
{  "mcpServers": {    "personamcp": {      "command": "uvx",      "args": ["personamcp", "--home", "/absolute/path/to/private/persona-data", "serve"]    }  }}

On Windows the command is C:\\absolute\\path\\.venv\\Scripts\\persona.exe. For Codex:

sh
codex mcp add personamcp -- /absolute/path/to/venv/bin/persona --home /absolute/path/to/private/persona-data servecodex mcp list

See Codex MCP setup. Other clients may use different configuration locations but need the same executable and arguments. Hosted clients that only support remote HTTP MCP cannot directly launch this local stdio server. PersonaMCP does not include a tunnel/HTTP bridge; exposing sensitive local data remotely requires a separate, explicit deployment decision.

With a prepared semantic model, allow up to 90 seconds for startup and 120 seconds for tools on slower machines. The native numerical runtime is loaded before stdio reader threads start to avoid Windows BLAS loader deadlocks. Weights are loaded on the first semantic query and cached. For Codex these settings belong in the server's configuration table:

toml
[mcp_servers.personamcp]command = "/absolute/path/to/venv/bin/persona"args = ["--home", "/absolute/path/to/private/persona-data", "serve"]startup_timeout_sec = 90tool_timeout_sec = 120

Tools:

ToolResult
get_style_profile(context?)Measured global or assigned-context profile
get_person_style(person)Communication statistics and evidence count for an exact person
search_messages(query, person?, platform?, limit?)Bounded outgoing text matches
find_similar_interactions(message, person?, context?, platform?, limit?)Similar real incoming/reply pairs
get_writing_context(message, person?, context?, platform?)Appropriate style and up to three quoted examples
get_persona_summary()Compact profile and database counts

Historical content is returned inside historical_quote, with an explicit untrusted-data notice. Tools supply evidence. Your agent generates the final reply and remains responsible for treating historical instructions as data and for deciding what to send to its model provider.

Skill setup

The reusable skill is skills/write-like-me/SKILL.md. It is also included in the wheel; persona skill-path prints its installed location. Copy its folder to the skill directory your agent discovers. For current Codex repository discovery:

sh
mkdir -p .agents/skillscp -R skills/write-like-me .agents/skills/

Windows PowerShell: New-Item -ItemType Directory -Force .agents/skills followed by Copy-Item -Recurse skills/write-like-me .agents/skills/. See Codex skill discovery.

Then ask the connected agent to use write-like-me, for example:

Write a short reply like me to Alex about “mai vii azi la cafea?”. Use PersonaMCP evidence.

The skill preserves supported casing, spelling, vocabulary, and length without forcing typos or copying old facts. It can also use persona writing-context through a local shell, or an explicitly supplied profile when MCP is unavailable. No global client configuration is changed by installation.

Privacy and data controls

The default data directory comes from your operating system's application-data location. --home /private/path or PERSONAMCP_HOME selects another location. Configuration is a local config.json; the database uses SQLite foreign keys, versioned schema, and transactional imports.

sh
persona stats --jsonpersona export-profile ./my-style.mdpersona export-profile ./my-style.json --jsonpersona delete-person "Name"persona delete-conversation CONVERSATION_IDpersona reset

Deletion/reset ask for confirmation; --yes is available for intentional scripting. Deleting a person removes entire conversations, including groups containing that person, to avoid keeping context about them. Profiles, FTS rows, and vectors are invalidated/purged and SQLite is vacuumed. reset also clears configured owner identities but retains downloaded, non-personal model assets. Original export files and any copied profiles remain in your control and are not deleted.

SQLite is not encrypted. Use a private data directory and disk encryption. Generated vocabulary and examples are sensitive too. If your agent uses an external LLM, tool results may reach that provider; “local-first” describes storage and computation, not the connected client's behavior. Read SECURITY.md for deletion and prompt-injection limits.

Offline benchmark

sh
persona benchmark --limit 30

The last 20% of interactions by timestamp form a holdout. Retrieval and style profiles use only earlier data. The held-out actual reply is used only for scoring. The default candidate is a retrieved historical reply, not an LLM-generated answer.

The report includes retrieval coverage, response-length similarity, vocabulary overlap, punctuation similarity, capitalization similarity, and their mean Style Similarity Score. When a local model is available, embedding similarity is reported separately. No raw benchmark replies are printed. CandidateResponseProvider is the extension point for a future explicitly configured generator. Scores do not prove identity imitation, authorship, relevance, or generation quality. Small or temporally uniform datasets cannot support this evaluation.

Architecture

text
exports → platform adapters → normalized conversations/messages                                  ↓                      SQLite + participants + FTS5                                  ↓                   bounded incoming/outgoing interactions                         ↙                    ↘              deterministic profiles      local embedding vectors                         ↘                    ↙                        retrieval/service layer                            ↙           ↘                         CLI          MCP stdio → writing skill → agent

The package uses a src/personamcp layout: importers, models, configuration, storage, analysis, embeddings, retrieval, shared service, MCP server, CLI, and benchmark. No external database or LLM provider is required. Schema compatibility is tracked with PRAGMA user_version; older versions reject a newer database instead of guessing how to read it.

Development, roadmap, and limits

Run uv sync --locked, then uv run pytest, uv run ruff check src tests, uv run ruff format --check src tests, and uv run mypy src. CI checks Windows/Linux and Python 3.11–3.13, with no personal data or model download. Public fixtures are synthetic. See CONTRIBUTING.md, CODE_OF_CONDUCT.md, and implementation decisions.

Next steps are additional export adapters (Discord, Telegram, Messenger, iMessage, Signal), explicit redaction rules, a faster local vector index for very large archives, richer language signals, and optional generation-based evaluation. They are not implemented in this MVP.

This version is a CLI/MCP engine. Export schemas can change; group recipient inference, psychological profiling, automatic typo correction, speech/media analysis, encryption at rest, cloud embedding providers, HTTP hosting, and a frontend are outside its current support.

MIT licensed © Robert Vind-Gardoș (@robyroro). Model weights and dependencies retain their own licenses; see the multilingual MiniLM model card and Sentence Transformers documentation.

來源:README.md,提交 52d546d

工具

0
工具後設資料尚未被收錄。

版本歷史

1
  1. v0.1.0最新Oct 1, 2026