Image, speech, audio, music and video generation and processing.
Autonomous data infrastructure that turns strategic objectives into verified datasets and live intelligence streams.
Enterprise voice AI platform for building, running and monitoring AI phone agents that hold inbound and outbound conversations at scale.
Exposes Cartesia's text-to-speech, speech-to-text and voice management APIs through MCP for use with AI agents and code editors.
API that handles the full speech pipeline (listening, thinking and speaking) so developers can build conversational voice agents without managing audio processing.
Official Deepgram MCP server giving AI editors access to Deepgram speech, transcription and audio-intelligence tools; the tool list is fetched from Deepgram's API at runtime.
ElevenLabs' platform for conversational voice agents; agents can be connected to external MCP servers to access tools.
Hosted MCP server that connects an ElevenLabs workspace to AI assistants: create and manage voice agents, generate speech, music, images and video, and estimate usage. OAuth, no local install.
MCP server to extract image metadata (EXIF, XMP, etc).
Zero-dependency browser video editor that AI agents can drive — JSON timeline, MCP + REST, live-reloading UI.
MCP server for Final Cut Pro XML. Lets Claude read, edit and generate real timelines: cut detection, markers, roles, transcript-based editing. Published on PyPI as fcp-mcp-server.
Using ffmpeg command line to achieve an mcp server, can be very convenient, through the dialogue to achieve the local video search, tailoring, stitching, playback,clip, overlay, concat and other.
Generate images, video, and audio with Glif's media-generation agent.
Creative agent that turns a prompt into a production-grade video, handling scriptwriting, image selection, voiceover and editing automatically.
Lets AI assistants use Hume's Octave text-to-speech from MCP clients: synthesize and play speech, and list, save and delete voices.
An MCP server providing tools for image processing operations.
Guardrailed video editing MCP server for AI agents. FFmpeg, Hyperframes, repurposing tools, Python client, and CLI. Local, fast, free.
MCP Server that uses the open weight Kokoro TTS models to convert text-to-speech. Can convert text to MP3 on a local driver or auto-upload to an S3 bucket.
Official Luma AI MCP server for creating images and videos with the Photon (image) and Ray (video) models from MCP clients.
This is an MCP server that allows you to directly download transcripts of YouTube videos.
MCP server that turns any video — YouTube, Instagram, TikTok, Loom, X, Vimeo, direct URLs, local files — into transcripts, key frames, OCR text, and metadata for AI agents.
Community MCP server that lets Claude and other clients use Stability AI's image generation and manipulation APIs.
Official MiniMax MCP server giving agents access to MiniMax text-to-speech, voice cloning, and video and image generation APIs.
Agent-native media processing via Transloadit's 86+ Robots: video encoding (HLS, H.264, VP9), image manipulation (resize, watermark, smart crop), document conversion, OCR, speech transcription, and.
AI Pocket AI co-editor for video montage — AI video editing plugin & MCP server for Claude Code, Codex, Hermes & OpenCode.
MCP server for OpenRouter — chat with 300+ LLMs (Claude, Gemini, GPT), analyze images / audio / video, generate images / speech / music / video (Veo 3.1, Sora, Seedance, Wan) from Claude Desktop.
An Open-Source Multimodal AIGC Solution based on ComfyUI + MCP + LLM.
AI-powered music production and post-production in REAPER via MCP — 172 tools for composition, mixing, mastering, batch editing, audio QC, and ReaScript automation.
Hosted MCP server for Replicate's HTTP API so agents can discover, compare and run thousands of image, video and audio models from clients like Claude, Cursor and VS Code.
Voice agent platform for call centers whose agents handle inbound and outbound calls, book appointments, navigate IVRs, transfer to humans and update CRMs.
MCP server for Suno AI music generation, lyrics, and cover workflows via Ace Data Cloud.
End-to-end voice AI platform with in-house telephony for deploying AI voice agents that automate phone calls for enterprises.
Recordings in. Agent-ready evidence out. Local-first MCP for transcripts, keyframes, OCR, speakers, and wall-clock evidence.
AI research lab building PALs: real-time conversational video AI agents that can see, hear, understand and respond in conversation.
Bridge between TwelveLabs' video understanding platform and AI assistants: video management, indexing, search, analysis and embeddings exposed as MCP tools.
Real-time voice AI infrastructure layer built on speech-native models that power fast, natural voice agents while preserving tone, cadence and pitch.
Descript's AI video editor agent that edits video, writes scripts and handles editing tasks from natural-language instructions inside Descript.
Platform for building AI voice assistants. Its MCP integration lets a Vapi assistant dynamically access tools from MCP servers during calls.
FFmpeg-powered MCP server for basic video and audio editing operations driven by AI assistants.
Add, Analyze, Search, and Generate Video Edits from your Video Jungle Collection.
Open-source library for building real-time, streaming voice applications and phone agents powered by large language models.
Natural voice conversations with Claude Code.
MCP server for YouTube that lets language models interact with YouTube content such as videos, transcripts, channels and playlists through a standard interface.
The AI-powered toolkit that grows your YouTube channel on autopilot.