Typing up an hour of audio by hand takes four to six hours. AI transcription does it in minutes, with timestamps, and gives you subtitle files you can upload straight to YouTube or your video editor.
This guide shows how to transcribe audio and video with Pynokio Transcribe, how to get the most accurate result, and which export format to choose.
What you can transcribe
- Audio: MP3, WAV, M4A, FLAC, OGG, OPUS and WebM.
- Video: MP4, MOV and WebM - the speech is extracted from the video for you.
- Size and length: files up to 64 MB and up to 2 hours long.
Interviews, podcasts, lectures, meetings, voice notes, YouTube videos and webinars all work well.
Step by step
- Open Transcribe and add your audio or video file.
- Choose the task. Transcribe keeps the original language. Translate turns the speech into English text - useful when the recording is in another language and you need English subtitles.
- Start transcription. The transcript appears with timestamps for each segment.
- Review it. Play the recording and skim for names, brands and technical terms - these are where errors usually hide.
- Export as TXT, SRT or VTT.
TXT, SRT or VTT?
| Format | What it contains | Use it for |
|---|---|---|
| TXT | Plain text without timestamps | Articles, show notes, quotes, search and AI summaries |
| SRT | Numbered subtitle blocks with start and end times | YouTube, most video editors and social platforms |
| VTT (WebVTT) | Subtitle cues in the web standard format | HTML5 video players, websites and online course platforms |
If you are unsure, download SRT. It is the most widely supported subtitle format.
How to get a more accurate transcript
- Start with clean audio. Background music, echo and people talking over each other are the main causes of errors.
- Use the original file. A recording that has been compressed several times loses detail that helps recognition.
- Trim long silences and music intros before uploading, if you can.
- Check proper nouns. Search the transcript for names and brands and fix them once with find-and-replace.
Tip: For video subtitles, keep each line short. If a subtitle block feels too long to read, split it in your editor - viewers read at roughly two to three words per second.
How AI transcription works
A speech recognition model listens to the audio in short windows and predicts the most likely words, using the surrounding context to decide between similar-sounding phrases. It also detects the language and marks where each segment starts and ends - those timestamps are what turn a transcript into subtitles.
Because the model relies on context, it handles everyday speech, accents and normal speaking speed well. It struggles most where people do too: overlapping voices, loud music, very distant microphones, and rare names it has never heard. That is why a quick human review of names and key terms is still worth the two minutes it takes.
Transcribe or translate?
Use Transcribe when you want the words exactly as spoken, in the original language - for captions, quotes, show notes and editing. Use Translate when the recording is in another language and you need to understand it or subtitle it in English. If you need subtitles in both languages, run the file twice and export both SRT files.
What to do with a transcript
- Subtitles and accessibility. Captions make videos usable for deaf and hard-of-hearing viewers and for the many people who watch with the sound off.
- Search visibility. A transcript or summary published next to a video or podcast gives search engines text to index.
- Repurposing. Turn an interview into a blog post, a newsletter or a set of social quotes.
- Localization. A transcript is the starting point for translated subtitles and dubbed voiceovers - see how to localize videos with AI.
- Voiceover remakes. Clean up a rough recording's script and regenerate it with AI text to speech.
Privacy
Files you upload are used to create your transcript, stay private to your account and are never published. Only upload recordings you have the right to process, and let people know when a conversation is being recorded and transcribed.
FAQ
Can I transcribe a video file directly?
Yes. Upload an MP4, MOV or WebM video and Pynokio extracts the speech for transcription.
How long can my recording be?
Up to 2 hours and 64 MB per file. Split longer recordings into parts.
Can it translate speech into English?
Yes. Choose Translate instead of Transcribe and you get English text from speech in another language.
What is the difference between SRT and VTT?
Both are subtitle formats with timestamps. SRT is the most widely supported by video platforms and editors; VTT is the web standard used by HTML5 video players.
Does the transcript include timestamps?
Yes. Every segment has timestamps, which are used for the SRT and VTT subtitle files.
Stop typing transcripts by hand. Upload a file and get text and subtitles in minutes.



