MinifyPic

Speech to Text

Talk and watch your words appear as text — dictation straight in the browser, with no app to install.

FreeLiveNo install
Loading tool…

How it works

A compact workflow from input to download.

1

Allow microphone access

Your browser will ask permission to use the microphone.

2

Start talking

Speak naturally and the transcription appears live as you go.

3

Edit and copy

Fix anything it misheard, then copy the text out.

Frequently asked questions

Which browsers support this?
Chrome and Edge have the best support for the Web Speech API's recognition side. Firefox and Safari support is more limited and varies by version, so if nothing happens when you start speaking, try Chrome.
Is my voice sent to a server?
Be aware that speech recognition, unlike speech synthesis, is generally processed in the cloud by the browser's own service rather than locally on your device. This is a genuine difference from the other tools here, and it means you should not dictate confidential material.
How do I get better accuracy?
Speak at a normal, even pace rather than slowly or loudly — recognition engines are trained on natural speech and exaggerated diction actually makes them worse. Use a decent microphone or a headset, reduce background noise, and dictate punctuation explicitly by saying 'comma' and 'full stop' where the engine supports it.
Can it handle accents?
Increasingly well, though performance still varies noticeably by accent, and speakers of less well-represented accents get measurably worse results. Selecting the closest matching language and regional variant in the settings helps considerably.

Dictation is faster than typing, and different

Most people speak at 120 to 150 words per minute and type at 40 to 60, so dictation is nominally two to three times faster — and yet many people who try it abandon it quickly. The reason is that dictating is a genuinely different cognitive activity from writing. Typing gives you constant opportunity to pause, reread, reconsider and revise, and that editing loop is woven into how most people compose. Speaking is linear and unforgiving; you must hold the shape of the sentence in your head before you begin it. The people who dictate well have learned to separate the two phases — get the words out fast and messy, then edit ruthlessly afterwards — rather than trying to compose polished prose in real time.

The privacy line worth knowing

There is an asymmetry in browser speech APIs that is not widely understood and genuinely matters. Speech synthesis — turning text into audio — runs locally on your device, so the text you have it read never leaves your machine. Speech recognition typically does not: in most browsers, the audio is sent to a cloud service for processing, because accurate recognition needs models far larger than a browser wants to ship. So dictating is not a private act in the way that having text read aloud is. Treat anything you dictate as though you had typed it into a search engine, and keep confidential material — medical details, legal matters, credentials — off the microphone.