Speech to Text — Live Dictation or File Transcription, Free
Dictate live with your browser's speech recognition, or transcribe audio and video files with Whisper on our server (free Google sign-in). Export text, SRT or VTT.
About Speech to Text
A speech-to-text tool turns spoken words into written text for dictation, meeting notes, voice memos, interviews and accessibility. The ZTools Speech to Text has two modes. Live microphone uses your browser's built-in Web Speech API (Chrome, Edge, Safari) and shows the text as you speak; in Chrome and Edge the audio is sent to the browser maker's speech service. Transcribe a file sends an audio or video recording (up to 15 minutes and 100 MB) to our own processing server, where the Whisper speech model detects the language, transcribes it or translates it into English, and returns text with timestamps that you can download as TXT, SRT or VTT. The file mode needs a free Google sign-in, and the server deletes the recording right after processing.
Use cases
- Dictation drafting. Speaking is ~150 wpm; typing is ~40-60 wpm. Long-form drafts (essays, blog posts, emails) go faster when you speak first and edit the transcript second.
- Meeting notes (live). During a one-on-one or interview, dictate key points instead of typing — eyes stay on the speaker, attention stays with the conversation.
- Interview and lecture transcripts. Record an interview, lecture or voice note, then transcribe the file with timestamps and download subtitles for the video.
- Accessibility. Users with motor impairments, RSI, or visual impairment can produce written text by speaking. Faster and less painful than alternative input methods.
How it works
- Pick a mode. "Live microphone" types as you speak. "Transcribe a file" turns a recording into text on our server.
- Live: allow the microphone and pick the language. The browser asks once. Choose the language you will speak, press start, and the text appears as you talk.
- File: sign in and add a recording. MP3, WAV, OGG, Opus, FLAC, AAC, M4A or a video (MP4, MOV, MKV, WebM, AVI), up to 100 MB and 15 minutes.
- File: choose the language and result. Let Whisper detect the spoken language or pick it, and choose text in that language or an English translation.
- Copy or download. Copy the text, or download a file transcript as timestamped text, SRT or VTT subtitles.
Examples
Input: 5 minutes of live dictation in English
Output: Several hundred words; clear speech in a quiet room needs few corrections.
Input: Spanish lecture recording (file mode)
Output: Spanish transcript with timestamps, or an English translation of it; download as SRT to caption the video.
Input: Technical content with many proper nouns
Output: Jargon and names are the most common errors in both modes; proofread them.
Frequently asked questions
Which browsers support live mode?
Chrome, Edge and Safari implement the Web Speech API; Firefox does not. The file mode works in any modern browser.
Is my voice uploaded?
In live mode it depends on the browser: Chrome and Edge send the audio to their makers' speech services, while Safari may recognise speech on the device. In file mode the recording goes over HTTPS to our own processing server, which deletes it right after processing.
How accurate is it?
Clear speech in a quiet room transcribes well in both modes. Accents, jargon, background noise and several people talking at once lower accuracy, so review the text before using it.
Why does live dictation sometimes stop?
Browsers end a live recognition session after a pause or a time limit. Press start again to continue; for long recordings, record first and use the file mode.
How long does file transcription take?
It depends on the length of the recording and how busy the server is; it can take about as long as the recording itself. Files can be up to 15 minutes.
Does it work offline?
Live mode works offline only where the browser recognises speech on the device. The file mode needs an internet connection to reach our server.
Pro tips
- Speak clearly at normal pace — too fast or too quiet drops accuracy noticeably.
- Use a real microphone for long sessions — built-in laptop mics introduce noise that degrades recognition.
- For long sessions, record the audio and use the file mode instead of dictating live.
- In file mode, pick the spoken language yourself when a recording starts with music or silence; automatic detection listens to the opening seconds.
- Always proofread before publishing — homophones (their/there/they're) consistently slip through.
Reviewed by Ahsan Mahmood · Last updated 2026-09-11 · Part of ZTools.
For the full,
formatted version of this page, please enable JavaScript and reload
https://ztools.zaions.com/speech-to-text.