terminal-use

io.github.computer-agent-labsv0.1.1更新于 Oct 9, 2026

Lets AI agents drive a real terminal: type, press keys, click, read the screen, take screenshots.

概览

AI 生成的概览

让 AI 助手获得一个真实终端会话,可以输入、按键、读取屏幕文本并截图,支持 vim、htop、REPL 等交互式程序。

功能
在真实伪终端上运行 shell,并由无头终端模拟器渲染,因此全屏和交互式程序可以正常工作,助手看到的是渲染后的屏幕而不是原始转义码。工具涵盖创建、列出和销毁会话;输入文本、按键、点击和滚动;等待命令结束、等待正则匹配或等待输出安静;以文本读取屏幕或回滚缓冲;以及把屏幕渲染为 PNG。会话可以运行交互式 shell 或单个命令,批量工具可在一次调用中发送多个输入。
适用场景
当助手需要操作普通“执行命令并返回输出”的 shell 工具无法处理的交互式程序时使用:vim 等编辑器、htop、REPL、会提问的安装程序、交互式 rebase、SSH 会话,或对 CLI/TUI 做端到端测试。也适合你想旁观并亲自在同一个会话中输入的场景。
运行要求
通过 npx 从 npm 包 terminal-use 启动的本地进程;需要 Node.js 20.19 或更高版本,支持 macOS、Linux 或 Windows。未声明任何账号、API 密钥或环境变量。仓库提供 Dockerfile,可在容器中运行,此时助手得到的是容器内的 shell 而非你机器上的 shell。
安装前请注意
它把以你的身份运行、没有沙箱的 shell 交给所连接的客户端,因此只应连接你愿意把终端交给它的客户端,并利用客户端的工具审批设置。它会通过所执行的命令写入和修改数据,在全屏程序中点击可能触发破坏性操作,因此点击预览默认开启。附加套接字按用户隔离并受限,但任何能连接的人都会获得与你相同的 shell。--login 标志会执行你的 shell 配置文件中的内容。

安装

在 SourceWeft 中

  1. 打开 控制台中的 terminal-use,将其添加到工作区。
  2. 为需要使用其工具的对话启用该服务。

Desktop only,通过 STDIO。 STDIO 服务会启动本地进程,因此需要 SourceWeft 桌面宿主。

其他 MCP 客户端

参照 仓库 中的启动说明。

README

terminal-use

An MCP server that lets an AI agent use a real terminal the way a person does: type, press keys, click, read the screen, take a screenshot.

[An agent opening vim, pasting a program, saving it, running it, and taking a screenshot — each step labelled with the tool call that made it]

Every frame above was drawn by terminal-use's own screenshot renderer; the caption on each is the tool call that produced it.

Most agent shell tools run a command and hand back its output. That breaks down for anything interactive — vim, htop, a REPL, an installer asking questions, git rebase -i, an SSH session. terminal-use gives the agent a shell on a real pseudo-terminal, rendered by a real terminal emulator, so full-screen and interactive programs work and the agent sees what you would see.

  • Real PTY, real emulator — colors, cursor movement and the alternate screen are interpreted, not passed along as escape codes.
  • Text and pixels — read the screen as plain text, or as a PNG when layout and color matter.
  • Waits properly — block until a command actually finishes, or until some text appears.
  • Watch along — attach your own terminal to any session and type alongside the agent.

Quick start

Requires Node.js 20.19 or newer, on macOS, Linux or Windows.

Claude Code

bash
claude mcp add terminal-use --scope user -- npx -y terminal-use

Other MCP clients (Claude Desktop, Cursor, and anything else that takes a command):

json
{  "mcpServers": {    "terminal-use": {      "command": "npx",      "args": ["-y", "terminal-use"]    }  }}

Start a new session in your client and ask it to do something in a terminal, for example: "Open vim, write a haiku into /tmp/haiku.txt, save and quit, then show me a screenshot of cat-ing it."

Server options go after the command: npx -y terminal-use --cols 100 --rows 40.

OptionDefault
--shell <path>$SHELL or /bin/bash; PowerShell on WindowsShell to run in new sessions
--cwd <path>where the server was startedWorking directory for new sessions
--cols <n> / --rows <n>120 / 30Terminal size
--scrollback <n>5000Lines of history kept
--loginoffStart shells as login shells (see below)

"command not found" inside a session

If programs that work in your own terminal (node, brew, pyenv…) are missing inside a session, the server probably inherited a bare environment. That happens when the MCP client is started from the Dock or a launcher instead of a terminal: your PATH is set up by your shell's profile files (~/.zprofile, ~/.bash_profile, ~/.profile), and nothing has read them.

A login shell reads those files. Turn it on for every session with the --login server flag, or for one session with login: true on terminal_create. It is off by default because it makes each shell slower to start and runs whatever your profile runs. It has no effect in PowerShell or cmd.exe, which have no login mode.

Tools

Every tool except terminal_create and terminal_list takes the sessionId that terminal_create returns.

ToolWhat it does
terminal_createStart a session: an interactive shell, or one program with command. Optional label, cols, rows, shell, cwd, env, login, scrollback, theme.
terminal_listList sessions.
terminal_destroyEnd a session.
terminal_typeType text. \n presses Enter. paste: true sends it as one paste.
terminal_pressPress a key or combination: Enter, Ctrl+C, ArrowUp, Shift+Tab, Alt+Enter, F5…
terminal_waitWait for the running command to finish, for a regex to appear, or for output to go quiet.
terminal_readRead the screen or scrollback as text.
terminal_screenshotRender the screen as a PNG.
terminal_clickLeft-click a cell, with a preview step.
terminal_scrollTurn the mouse wheel over a cell.
terminal_batchSend several inputs in one call and get the screen back once.
terminal_resizeChange the terminal size.
terminal_resetClear the screen and scrollback, or restart the shell with hardReset: true.

Docker

The repository has a Dockerfile for running the server in a container:

bash
docker build -t terminal-use .docker run -i --rm -v "$PWD":/workspace terminal-use

In an MCP client, use docker as the command with run -i --rm terminal-use as its arguments. The terminals the agent gets are then shells inside the container, not on your machine: it sees only what you mount, which is the point if you want it boxed in. Attaching from your own terminal is not available in this setup.

The skill

terminal-use comes with an Agent Skill: a short guide for the agent on how to use these tools well — when to use a command session, how to wait, how to read what is selected — plus recipes for vim, pagers, REPLs, menu-driven programs, ssh prompts and end-to-end testing of a TUI. It lives in skills/terminal-use.

  • Hosts that load skills from MCP servers get it automatically. The server implements the MCP Skills extension (io.modelcontextprotocol/skills): the skill is listed by skills/list and its files are served as skill://terminal-use/... resources. Few hosts support this yet.

  • Hosts that load skills from disk can install the same files. For Claude Code:

    bash
    cp -r "$(npx -y terminal-use skills-dir)/terminal-use" ~/.claude/skills/

The skill is optional. Without it the agent still has the tool descriptions and the server's built-in instructions.

Shell sessions and command sessions

By default a session is an interactive shell: type commands into it as you would at a prompt.

Pass command to terminal_create to run one program in the terminal instead:

json
{"command": "npm test -- --watch", "cwd": "/path/to/project", "env": {"CI": "1"}}

The command goes through the session's shell (sh -c, PowerShell -Command, or cmd /c), so quoting, pipes and redirection work as they do at that shell's prompt. Input goes straight to the program. When it exits:

  • its exit status is reported (by the call that was in progress, by terminal_wait, and in terminal_list);
  • the final screen stays readable with terminal_read and terminal_screenshot;
  • input tools return an error with the exit status rather than restarting anything;
  • terminal_reset with hardReset: true runs it again, and terminal_destroy removes it.

This is the mode for testing a CLI or TUI: launch it, drive it, check how it ended.

Watching and typing along

terminal_create returns a command you can run in your own terminal to join the session:

node /path/to/terminal-use/bin/terminal-use.js attach 3 --socket /tmp/terminal-use-501/41234-3.sock

You see what the agent sees and can type into the same shell. It works like a shared tmux session: several people can attach at once, and Ctrl+] detaches.

  • You get the current screen and recent scrollback on connect, not a blank terminal.
  • --resize makes the session follow your window size. By default the session keeps its own.
  • --socket picks the server. Each MCP client runs its own terminal-use, and they all number sessions from 1; without --socket, attach <id> works when only one running server has that id and lists the candidates otherwise.
  • On Windows the session is reached through a named pipe instead of a socket file; the command you are given works the same way.
  • Sockets are per-user (0600, inside a 0700 directory under the system temp dir). Anyone who can connect gets a shell as you, so they are not exposed any further than that.

How it works

agent ──MCP──▶ terminal-use ──▶ node-pty ──▶ your shell                  │                  └─ @xterm/headless  ◀── everything the shell prints                        │                        ├─ terminal_read        (text)                        └─ terminal_screenshot  (PNG)

node-pty runs the shell on a pseudo-terminal. Everything it prints is fed to a headless xterm.js emulator, and that emulator's buffer is the single source of truth: reads and screenshots both come from it, so the agent gets the rendered screen rather than a stream of escape codes.

Details

Waiting for things

terminal_type, terminal_press and terminal_click return once output has been quiet for a moment (idleMs, default 200 ms) and never wait longer than 10 seconds. That suits keystrokes. It doesn't suit a build, which can be silent for a while long before it is done. For anything slow, follow up with terminal_wait:

  • Default — wait for the command to finish. terminal-use asks the kernel which process owns the terminal's foreground. A shell hands the terminal to each command it runs and takes it back afterwards, so when the shell owns it again, the prompt is back. This needs no shell integration or prompt parsing, and works for commands that print nothing.
  • pattern — wait for a regex to match the screen. For things that never exit (Listening on port), REPL prompts, or a particular state of a TUI. ^ and $ match at line starts and ends; a leading (?i) or (?s) sets further flags.
  • until: "quiet" — wait for output to stop for quietMs (default 1 s). The same rule the typing tools use, without the 10-second cap.

In a command session, the default mode waits for the program to exit and reports its exit status.

timeoutMs defaults to 30 seconds (maximum 10 minutes). A timeout isn't an error: the response says the command is still running, and you can wait again.

On Windows there is no way to ask who owns the terminal, so in a shell session the default mode waits for two seconds of silence instead, and says it is a guess. Prefer a pattern there, or run the program as a command session, where waiting for it to exit works on every platform.

Limits worth knowing: background jobs (cmd &) don't count as running. Inside a nested program such as ssh or a REPL, the outer shell doesn't get the foreground back until that program exits, so use pattern or until: "quiet" there. No exit status is reported; run echo $?.

Reading the screen

terminal_read returns text in screen-sized pages: page: 0 is the current screen, page: 1 the one before it, and so on back through scrollback.

Lines the terminal wrapped at its right edge are joined back into the single line the program printed (joinWrapped: false gives one line per screen row).

Text can't show color, so it can't show which menu entry is selected. Reads therefore end with a list of what is highlighted on screen — text in reverse video or on a background color — with the row and columns terminal_click takes:

Highlighted on screen (reverse video or background color; screen row, columns):  row 7, cols 3-18: "Unstaged changes"

If most of the screen is colored panels, the list is replaced by a pointer to terminal_screenshot. Turn it off with highlights: false.

The cursor is marked with ▌. On an empty cell it simply takes the place of the blank. On a character it is inserted in front of that character, which shifts the rest of that one line right by a column; the header says which character it is on. The marker is left out while the program hides the cursor (most full-screen programs do), and cursor: false returns the text untouched.

Screenshots

terminal_screenshot renders the screen with the bundled JetBrains Mono. This is what a call returns, here with a visual selection in vim:

[A terminal_screenshot of vim with Python syntax highlighting and four lines selected]

Sessions take a theme at creation: dark (default), light, solarized-dark or solarized-light.

JetBrains Mono covers Latin, Greek, Cyrillic, box-drawing and common symbols. Emoji, Chinese/Japanese and Korean text fall back to fonts already on the system, because bundling them would add tens of megabytes:

EmojiCJKHangul
macOSApple Color EmojiHiragino Sans GB / PingFangApple SD Gothic Neo
WindowsSegoe UI EmojiMicrosoft YaHei / Yu GothicMalgun Gothic
LinuxNoto Color EmojiNoto Sans CJKNoto Sans CJK

macOS has these out of the box, and Windows has the emoji font (the East Asian fonts come with the matching language features). On Debian or Ubuntu, install them with apt install fonts-noto-color-emoji fonts-noto-cjk. Without them (a bare Docker image, say) those characters render as empty boxes in screenshots. Nerd Font and Powerline icons render as boxes everywhere. terminal_read is unaffected and always returns the real characters.

Typing, pasting and batching

terminal_type sends characters as keystrokes. For multi-line text going into an editor, a REPL or a shell prompt, add paste: true: programs that support bracketed paste receive it as one paste and insert it verbatim, without auto-indenting or running each line as it arrives. Programs that don't support it get the plain characters.

terminal_batch sends a list of inputs in one call and returns the screen once at the end, which saves a round trip per keystroke when the steps are already known:

json
{  "sessionId": 1,  "actions": [    {"type": "press", "key": "ArrowDown", "count": 3},    {"type": "press", "key": "Enter"},    {"type": "wait", "pattern": "Commit message"},    {"type": "type", "text": "Fix typo"},    {"type": "press", "key": "Ctrl+S"}  ]}

Actions are type, paste, press, click, scroll and wait (a fixed ms, or a pattern to appear). The batch is checked before anything is sent, and stops at the first action that fails, reporting how far it got.

Clicking and scrolling

  • Left button only. Coordinates are 1-indexed; (1, 1) is the top-left cell.
  • preview is on by default. A preview returns a screenshot with a ring drawn around the target cell and sends no click. Repeat the call with preview: false to click. Full-screen programs often have destructive actions one click away, so it is worth the extra step.
  • terminal_scroll turns the wheel over a cell. Programs that track the mouse get wheel events there, so the pane under the pointer scrolls; full-screen programs that don't (less, man) get arrow keys, as in a normal terminal. At a shell prompt there is nothing to scroll — read earlier output with terminal_read and page.
  • A real click needs the program to have turned on mouse reporting — vim with set mouse=a, fzf, lazygit, htop and most modern TUIs do. At a plain shell prompt the call returns an error instead of printing escape codes into your command line.

Session lifecycle

  • Sessions are independent. Calls to one session run in order; calls to different sessions don't block each other.
  • If a shell session's shell exits, the session sits idle for six hours, or the 50-session limit is reached, the session is shut down but its id stays reserved for 30 days. The next call to that id starts a fresh shell with the same size, shell, working directory and theme, and returns a notice saying what happened — including the last screen the old shell printed, if it exited on its own. The command in that call is not run; send it again if you still want it.
  • Idle means no tool calls and no output. A dev server that is still printing is not idle.
  • terminal_destroy ends a session for good; its id is not reserved.
  • When the MCP client disconnects or the server is stopped, every shell is closed and every attach socket removed.

Protocol support

Built on the official MCP TypeScript SDK (v2). Over stdio it speaks both the 2026-07-28 revision of the protocol, which is stateless, and the earlier handshake-based revisions; the client's first message decides which.

Stateless refers to the protocol, not the terminals: sessions live in the server process and are addressed by the sessionId you pass on each call. The server also provides usage instructions, titles and behavior hints for each tool, cancellation, progress updates from terminal_wait, and the Skills extension described above.

Platform support

macOS (arm64, x64)Supported; tested in CI
Linux (arm64, x64)Supported; tested in CI on Node 20, 22 and 24
Windows (x64)Supported; tested in CI against PowerShell and cmd.exe

On Windows, sessions run through ConPTY, the default shell is Windows PowerShell, and two things differ: terminal_wait cannot detect that a shell command has finished (see Waiting for things), and login does nothing.

The native dependencies (node-pty, @napi-rs/canvas) ship prebuilt binaries, so no compiler is needed to install.

Development

bash
git clone https://github.com/computer-agent-labs/terminal-use.gitcd terminal-useyarn installyarn build      # compile src/ to dist/yarn test       # build, then run the full suite

Other scripts:

bash
yarn dev            # run the server straight from sourceyarn lint           # eslintyarn typecheck      # tsc, no outputyarn test:unit      # unit tests onlyyarn test:windows   # what CI runs on Windows: shell-independent tests + tests/windowsyarn smoke          # quick end-to-end check without an MCP clientyarn demo           # regenerate docs/demo.gif (needs ffmpeg, vim and python3)

To point your MCP client at a local checkout, build it and use the path to the bin script:

bash
claude mcp add terminal-use --scope user -- node "$PWD/bin/terminal-use.js"

The code is laid out by layer: src/pty (spawning, key and mouse encoding), src/emulator (the xterm buffer, rendering, waiting), src/session (one terminal), src/attach (the socket you attach through) and src/mcp (the tools). The agent skill is in skills/.

Security

terminal-use gives the connected client a shell on your machine, running as you, with no sandbox. Connect it only to clients you would trust with a terminal, and use your client's tool-approval settings to control what runs unprompted. It opens no network ports; attach sockets are restricted to your own user. See SECURITY.md for details and for how to report a vulnerability.

Contributing

Issues and pull requests are welcome. CONTRIBUTING.md covers setting up, running the tests, and what a good pull request looks like. This project follows the Contributor Covenant.

License

MIT. The bundled JetBrains Mono fonts are licensed under the SIL Open Font License.

来源:README.md,提交 2f5951b

工具

0
工具元数据尚未被收录。

版本历史

1
  1. v0.1.1最新Oct 9, 2026