Get started
Welcome to Pynokio
One workspace for AI voice, music, sound, transcription, images and video - built on our own audio engines.
What is Pynokio
Pynokio is a creative AI platform. You describe what you need, and Pynokio turns it into speech, a song, a sound effect, a transcript, an image or a video clip. Everything you make lands in one library, ready to download or reuse in your next project.
The audio side runs on engines we build and operate ourselves: speech synthesis, voice cloning and design, music, sound effects and transcription. For images and video, Pynokio connects to leading generative models through their APIs, so you can pick the right model for each job without juggling separate accounts.
What you can create
- Text to Speech
- Natural narration in 30 languages, streamed as it is generated, with a library of ready-made voices.
- Voice cloning and design
- Clone your own voice from a short sample, or design a new one from a written description.
- Music
- Full tracks from a description, with optional lyrics, instrumental mode, tempo and length.
- Sound effects
- Foley, ambiences and short effects from a text prompt, ready to drop into an edit.
- Transcribe
- Turn audio or video into text with timestamps, or translate the speech into English.
- Image and video
- Generate and edit visuals with leading image and video models, from a prompt or a reference.
- Studio
- Describe a project in plain words and let the Studio assistant plan it and run the right tools for you.
How it works
- Choose a tool. Open Text to Speech, Music, Sound effects, Transcribe, Image, Video or Studio.
- Describe or upload. Write a prompt, paste a script or add a reference file.
- Generate. Credits are reserved when a generation starts, only the actual cost is charged, and failed generations are refunded.
- Use the result. Download it, find it later in your library, or build on it in your next step.
Paid plans are monthly: at the start of every billing period your balance is topped up to the plan's credit allowance.
Built responsibly
Every generated voice carries an inaudible watermark, and generated files come with a machine-readable provenance record. Anyone can check an audio file for free with the Audio Detector. Voice cloning requires the consent of the voice owner, as set out in our Voice Policy.
Explore the docs
- ModelsModels available in your workspace
- AI TransparencyWhat is generated and how it is marked
- Voice PolicyConsent and rules for voice cloning
- Audio DetectorCheck a file for the Pynokio watermark
- Machine-readable markingWatermark and provenance records