Writing modes
“Writing mode” is a separate setting from which language you speak. Language is the input — what you’re saying. Writing mode is the output — what script the committed text ends up in. You find it on the Models screen, next to the language and model you’ve picked.
| Language | Native script | English output |
|---|---|---|
| Bengali | Yes | Yes |
| English | Yes | — |
| Gujarati | Yes | Yes |
| Hindi | Yes | Yes |
| Hinglish | Yes | — |
| Kannada | Yes | Yes |
| Malayalam | Yes | Yes |
| Marathi | Yes | Yes |
| Punjabi | Yes | Yes |
| Tamil | Yes | Yes |
| Telugu | Yes | Yes |
Native script — the default
Every one of the eleven languages writes in its own script by default: Devanagari for Hindi, Bengali for Bengali, Gurmukhi for Punjabi, and so on. This is what you get unless you deliberately change it, and it’s what every speech model in betterflo is built to produce.
Roman letters — not shipped, for any language
The writing-mode picker also lists a Roman-script option — the same Indic language, written in the English alphabet instead of its own script. It is selectable, and it does nothing today. No speech model in betterflo writes any Indic language in the Roman alphabet yet; picking it stores your preference and nothing else changes. An earlier attempt to fake it by transliterating the native-script output was tried and rejected — it produced results wrong often enough that shipping it would have been worse than not offering it. The setting turns itself on the day a model that actually does this ships; nothing about how you use betterflo changes if it never does.
English output — ships, with a real caveat
For 9 of the eleven languages, betterflo can translate what you said into English instead of writing it in its native script. It’s a real, shipping feature — not a preview — available on the two higher device tiers.
It runs as a separate step, after your speech is already transcribed and cleaned up: a dedicated translation model (229 MB, downloaded once, on top of your language’s own speech model) reads the finished transcript and produces an English version of it. It is not a built-in “translate” mode inside the speech model itself — an earlier version worked that way and was pulled because it tended to produce fluent English that had nothing to do with what was actually said.
That same failure mode is still the honest limit of English output today, just less often. In a blind evaluation, only about 50% of translated sentences were judged to actually carry the meaning of what was said (blind, paired, 40 Vaani clips in bn and te; self-consistency 6/6) — against a roughly 70% bar the team had set before shipping. The team shipped it anyway, as an explicit decision, because the alternative was withholding a real capability while a better fix was found, not because the number cleared the bar. Even with a perfect transcript handed to the same translator — the best this design can do — the same evaluation put the ceiling at 75% (the best any translator could do here — the same model fed perfect reference text).
What that means in practice: when English output is wrong, it doesn’t look wrong. It reads as ordinary, grammatical English — not garbled, not obviously mistranslated — and it can say something different from what you actually said. You can’t tell from the English text alone whether it’s one of the roughly half that’s right or one that isn’t. It also takes noticeably longer to appear than native script does, with nothing on screen telling you a translation is running while you wait.
If you’re dictating something where being wrong quietly matters — anything you won’t proofread against what you actually said — native script is the safer default. English output is genuinely useful for a quick gist or a low-stakes message; treat anything it produces as a first draft, not a transcript.
Related
- Pick your languages — which languages ship, and their native-script models.
- Troubleshooting — what to do if English output looks confidently wrong, or the translator won’t download.