AI Voice Director: Multi-Character TTS Synthesis Tool
Regular price
£57.99
Regular price
£57.99
Sale price
Unit price/ per
SAVE
Sold out
AI Voice Director: Multi-Character TTS Synthesis Tool
The AI Voice Director is a specialized skill designed to facilitate multi-character text-to-speech (TTS) synthesis for audio content production. It allows users to select distinct voices for different characters from a voice catalog, configure per-segment synthesis parameters, and manage the audio stitching process using planning and crossfading techniques.
What this skill does
Utilizes a voice catalog comprising Kokoro, DIA, and Qwen3-TTS families to match character profiles with appropriate voice options.
Allows setting of synthesis parameters such as speed and stability for individual segments of a script.
Plans and executes the stitching of audio segments using tools like ffmpeg, including segment rendering and crossfade transitions.
Integrates descriptively styled prompts to design voices on supported TTS models, enhancing the natural sound and feel of the synthesized speech.
Operates in conjunction with podcast-producer to read scripts and hands off the rendered audio plan to episode-publisher for final production.
Who it is for
This skill is intended for developers and production teams using AI coding agents such as Claude Code, Cursor, and Codex. It is particularly suited for those involved in creating audio content with multiple character roles and complex dialogue needs.
Use cases
Producing voiced content from pre-existing scripts that have passed the script linting process, ensuring coherent and professional audio output.
Creating audio dramas or podcasts where different characters require distinct vocal expressions and tonal qualities.
Generating audiobooks with diverse character dialogues, enhancing listener engagement through varied voice profiles.
Technical details
The AI Voice Director uses a voice catalog that maps character profiles to specific voices identified by an "ID" system. This mapping ensures character-to-voice congruence, which is critical for maintaining audio authenticity. The synthesis process leverages ffmpeg for audio assembly, which includes crossfading and segment concatenation. The skill operates independently of any particular TTS engine but requires confirmation of the chosen engine's availability based on the voice-catalog.md reference. The skill is optimally used in concert with podcast-producer for input and episode-publisher for audio output handling.
Source & Licence
This package is built on open-source work published by X33834 (X33834/awesome-skillkit) and distributed under Apache-2.0. The original licence text and copyright notice are included in your download.
Personal and commercial use, modification and redistribution are permitted under the Apache License 2.0, which also includes an express patent grant. Attribution and any NOTICE file must be retained.
Your purchase covers curation, licence verification, packaging, documentation and instant delivery. It does not grant exclusive rights to the underlying open-source code, which remains available under its original licence.
Delivery & Support
Delivery: instant — a secure download link is emailed to you as soon as payment is confirmed.
Format: ZIP archive containing the skill files, documentation and the original licence.
Updates: updates are included only where stated on this page.
Refunds
This is a digital product delivered immediately after purchase. By completing your order you request immediate delivery and acknowledge that, once the download has been accessed, the statutory right to cancel no longer applies to the extent permitted by law. Refund requests are handled in accordance with our published Refund Policy.
Claude, Codex, Gemini and Cursor are trademarks of their respective owners. MCP Cart is an independent marketplace and is not affiliated with, endorsed by, or sponsored by any of them. Compatibility references describe interoperability only.