Friskription → Blog → Live Caption
Chrome Live Caption vs translated subtitles
Live Caption is a genuinely good accessibility feature that a lot of people use as a translation tool. It is worth knowing precisely where that stops working, because the line is sharper than it looks.
The short answer
- Use Live Caption when the audio is in a language it supports, you want the gist, and you would rather nothing left your machine. It is free, on-device and already installed.
- Use a translating overlay when the spoken language is outside Live Caption's set, when names and terms have to stay consistent, or when you want subtitles on the video rather than in a floating box.
- The decisive difference is context: Live Caption translates with general-purpose machine translation, line by line. An overlay built on a language model reads each line against the ones before it.
- The other decisive difference is privacy, and it runs the other way: Live Caption processes everything locally and sends nothing anywhere.
- They are not mutually exclusive. Live Caption costs nothing to try first.
What Live Caption actually is
Live Caption is an accessibility feature. Chrome downloads a speech-recognition model to your machine, and from then on any audio playing in the browser is transcribed into a small floating caption box. It works offline, it works on anything that makes sound, and nothing is sent to a server. For someone who is deaf or hard of hearing, on content in a language they read, it is excellent and free.
Translation came later and is a smaller feature bolted to a captioning one. That is the source of nearly everything below: the parts people run into are the parts that were designed for accessibility rather than for reading a foreign-language stream.
Where the two differ
| Chrome Live Caption | A translating overlay | |
|---|---|---|
| Where it runs | On your machine, offline | Transcription and translation on a server, or your own API keys |
| Source languages | A limited set | Detected automatically, across a much wider range |
| Translation | General-purpose machine translation | A language model reading each line in context |
| Where the text appears | A fixed floating box | An overlay on the video — drag, resize, restyle |
| Consistent names | No mechanism for it | A glossary pins spellings for the session |
| Original language alongside | No | Optional second line under each cue |
| Speaker labels | No | Up to five voices |
| Keeping the text | No export | Transcript export; SRT and VTT from your own files |
| Cost | Free, unlimited | 30 minutes a week free, then one-time packs from $1.99 |
Why context is the one that matters
Line-by-line machine translation treats each caption as an independent string. That is fine for a written paragraph, where the sentences arrive whole. Live speech does not arrive whole: a sentence is split across two captions, a pronoun in the second refers to a noun in the first, and a language that drops subjects — Japanese and Korean do this constantly — leaves the crucial information in a caption that has already scrolled away.
A translator that reads each line against the previous ones can resolve that. It also stops the same name being transliterated three different ways in five minutes, which is the single most tiring failure mode of line-by-line translation on a stream full of proper nouns.
There is a measurable input to this too. The transcript sets the ceiling for everything after it: no translator can recover a word that was misheard. Friskription transcribes with ElevenLabs Scribe v2, which publishes a 3.3% word error rate on English benchmarks — against 5.3% for Deepgram Nova-3 and 7.7% for Whisper large-v3. Every engine scores worse on a noisy live stream than on a benchmark clip, but the ordering is what carries over.
Where Live Caption wins outright
Privacy, without qualification. Live Caption runs on your machine. No audio leaves it, no server is involved, and it works with no network at all. If you are captioning something confidential, that is not a feature you trade away lightly, and it is genuinely the better choice.
The nearest equivalent on the other side is connecting your own OpenAI, Anthropic or ElevenLabs API key: the audio then goes from your browser straight to the provider you already have a contract with, and never reaches Friskription's servers. That is a meaningfully different guarantee from on-device, and worth being clear about rather than blurring.
Live Caption is also free and unmetered. Thirty free minutes a week is a real allowance, and it is not unlimited. For hours of casual listening in a language Live Caption supports, it is the sensible tool.
Which to reach for
Live Caption, when…
- The audio is in a language it supports and you mainly want the gist.
- Nothing may leave the machine.
- You want captions for accessibility rather than translation.
- You are listening for hours and do not want to think about minutes.
A translating overlay, when…
- The spoken language is outside Live Caption's set — which covers most Japanese VTuber streams and most Korean streams.
- Names, characters and recurring terms have to stay spelled the same way.
- You want the subtitle on the video, where your eyes already are, rather than in a box in the corner.
- You want the original line underneath, for learning or for checking.
- Several people are talking and you need to know which is which.
- You want to keep the text afterwards.
Questions
Is Chrome Live Caption good enough for translating streams?
For the gist of a stream in a language it supports, often yes. It struggles where live translation is hardest: source languages outside its set, sentences split across captions, dropped subjects, and proper nouns that need to stay consistent. It is an accessibility feature with translation added, not a translation tool.
Does Live Caption work on Japanese or Korean streams?
Only if the spoken language is in the set Chrome supports for captioning, which is more limited than most people expect and is the usual reason people look for something else for VTuber and Korean streams. It is free, so it costs nothing to check on the stream you care about.
Is Live Caption free?
Yes, entirely, with no limit and no account. It is part of Chrome and runs on your machine.
Which is more private?
Live Caption, clearly — it processes everything on-device and sends nothing anywhere. The closest alternative is connecting your own API keys, which keeps audio between your browser and a provider you already have a contract with, but that is not the same as never leaving the machine.
Can I use both?
Yes. Try Live Caption first, since it is free and instant; reach for a translating overlay when the language is not covered, when the names keep changing, or when you want the subtitle on the video rather than in a corner box.
Why does Live Caption put the text in a separate box?
Because it was built as a system-wide accessibility feature that captions all browser audio, including audio with no video attached. An overlay drawn on the video is only possible for a tool that knows it is looking at a video player.
Does either of them need the stream to have captions?
No — that is what both have in common, and why both work on Twitch. Each transcribes the audio itself rather than reading a caption track, so neither needs anything from the platform or the streamer.
Try the other half of the comparison
Thirty free minutes every week. No card, no AI accounts.
Add it to Chrome