
Speech to Text (Whisper) for Audio and Video URLs
io.github.saulius876-lgtmv1.0.0Updated Oct 2, 2026
Whisper speech to text for podcast, audio and video URLs, 90+ languages, timestamped output.
Overview
AI-generated overview
Transcribes audio and video from URLs into timestamped text using Whisper, supporting over 90 languages.
- What it does
- This remote MCP server gives an assistant a speech-to-text capability: it takes a podcast, audio, or video URL and returns a transcript. Output is timestamped and covers more than 90 languages. It is hosted by Apify and reached over streamable HTTP, so no local package is installed.
- When to use it
- Use it when you want an assistant to turn a spoken-word link — a podcast episode, interview, or video — into searchable, timestamped text without running transcription software locally.
- Requirements
- A remote MCP client that supports streamable HTTP, pointed at the hosted endpoint. Network access is required. The manifest declares no authentication, environment variables, or headers, but the provider may still require an account or token at connection time.
Before you install
The audio or video URL you submit is sent to the third-party provider for processing, so avoid confidential or private media. No credentials are declared in the manifest, but confirm what the provider asks for before connecting, and note that long recordings may consume paid provider quota.
Installation
In SourceWeft
- Open Speech to Text (Whisper) for Audio and Video URLs in the dashboard and add it to a workspace.
- Enable the server for the chats that should use its tools.
Web executable via Streamable HTTP. Remote servers run from the web runtime once configured in a workspace.
Other MCP clients
Add this to your client's mcpServers config.
{
"mcpServers": {
"speech-to-text": {
"type": "http",
"url": "https://mcp.apify.com/?tools=sauliusautomatesit/media-transcriber"
}
}
}Tools
0Tool metadata has not been indexed yet.
Version history
1- v1.0.0LatestOct 2, 2026

