FriskriptionBlogLive Caption

Chrome Live Caption vs translated subtitles

Live Caption is a genuinely good accessibility feature that a lot of people use as a translation tool. It is worth knowing precisely where that stops working, because the line is sharper than it looks.

The short answer

  • Use Live Caption when the audio is in a language it supports, you want the gist, and you would rather nothing left your machine. It is free, on-device and already installed.
  • Use a translating overlay when the spoken language is outside Live Caption's set, when names and terms have to stay consistent, or when you want subtitles on the video rather than in a floating box.
  • The decisive difference is context: Live Caption translates with general-purpose machine translation, line by line. An overlay built on a language model reads each line against the ones before it.
  • The other decisive difference is privacy, and it runs the other way: Live Caption processes everything locally and sends nothing anywhere.
  • They are not mutually exclusive. Live Caption costs nothing to try first.

What Live Caption actually is

Live Caption is an accessibility feature. Chrome downloads a speech-recognition model to your machine, and from then on any audio playing in the browser is transcribed into a small floating caption box. It works offline, it works on anything that makes sound, and nothing is sent to a server. For someone who is deaf or hard of hearing, on content in a language they read, it is excellent and free.

Translation came later and is a smaller feature bolted to a captioning one. That is the source of nearly everything below: the parts people run into are the parts that were designed for accessibility rather than for reading a foreign-language stream.

Where the two differ

Chrome Live Caption A translating overlay
Where it runs On your machine, offline Transcription and translation on a server, or your own API keys
Source languages A limited set Detected automatically, across a much wider range
Translation General-purpose machine translation A language model reading each line in context
Where the text appears A fixed floating box An overlay on the video — drag, resize, restyle
Consistent names No mechanism for it A glossary pins spellings for the session
Original language alongside No Optional second line under each cue
Speaker labels No Up to five voices
Keeping the text No export Transcript export; SRT and VTT from your own files
Cost Free, unlimited 30 minutes a week free, then one-time packs from $1.99

Why context is the one that matters

Line-by-line machine translation treats each caption as an independent string. That is fine for a written paragraph, where the sentences arrive whole. Live speech does not arrive whole: a sentence is split across two captions, a pronoun in the second refers to a noun in the first, and a language that drops subjects — Japanese and Korean do this constantly — leaves the crucial information in a caption that has already scrolled away.

A translator that reads each line against the previous ones can resolve that. It also stops the same name being transliterated three different ways in five minutes, which is the single most tiring failure mode of line-by-line translation on a stream full of proper nouns.

There is a measurable input to this too. The transcript sets the ceiling for everything after it: no translator can recover a word that was misheard. Friskription transcribes with ElevenLabs Scribe v2, which publishes a 3.3% word error rate on English benchmarks — against 5.3% for Deepgram Nova-3 and 7.7% for Whisper large-v3. Every engine scores worse on a noisy live stream than on a benchmark clip, but the ordering is what carries over.

Where Live Caption wins outright

Privacy, without qualification. Live Caption runs on your machine. No audio leaves it, no server is involved, and it works with no network at all. If you are captioning something confidential, that is not a feature you trade away lightly, and it is genuinely the better choice.

The nearest equivalent on the other side is connecting your own OpenAI, Anthropic or ElevenLabs API key: the audio then goes from your browser straight to the provider you already have a contract with, and never reaches Friskription's servers. That is a meaningfully different guarantee from on-device, and worth being clear about rather than blurring.

Live Caption is also free and unmetered. Thirty free minutes a week is a real allowance, and it is not unlimited. For hours of casual listening in a language Live Caption supports, it is the sensible tool.

Which to reach for

Live Caption, when…

A translating overlay, when…

Questions

Is Chrome Live Caption good enough for translating streams?

For the gist of a stream in a language it supports, often yes. It struggles where live translation is hardest: source languages outside its set, sentences split across captions, dropped subjects, and proper nouns that need to stay consistent. It is an accessibility feature with translation added, not a translation tool.

Does Live Caption work on Japanese or Korean streams?

Only if the spoken language is in the set Chrome supports for captioning, which is more limited than most people expect and is the usual reason people look for something else for VTuber and Korean streams. It is free, so it costs nothing to check on the stream you care about.

Is Live Caption free?

Yes, entirely, with no limit and no account. It is part of Chrome and runs on your machine.

Which is more private?

Live Caption, clearly — it processes everything on-device and sends nothing anywhere. The closest alternative is connecting your own API keys, which keeps audio between your browser and a provider you already have a contract with, but that is not the same as never leaving the machine.

Can I use both?

Yes. Try Live Caption first, since it is free and instant; reach for a translating overlay when the language is not covered, when the names keep changing, or when you want the subtitle on the video rather than in a corner box.

Why does Live Caption put the text in a separate box?

Because it was built as a system-wide accessibility feature that captions all browser audio, including audio with no video attached. An overlay drawn on the video is only possible for a tool that knows it is looking at a video player.

Does either of them need the stream to have captions?

No — that is what both have in common, and why both work on Twitch. Each transcribes the audio itself rather than reading a caption track, so neither needs anything from the platform or the streamer.

Try the other half of the comparison

Thirty free minutes every week. No card, no AI accounts.

Add it to Chrome