← Writing

Hinglish is not a language, which is why your caption tool cannot write it

Caption tools now advertise language lists in the hundreds. Code-mixed speech is missing from all of them, and the reason is architectural rather than a gap in coverage.

LumaCaption · · 6 min read

Hinglish is absent from caption tools advertising over a hundred languages because a language selector asks which single language the audio is in, and code-mixed speech has no correct answer to that question — the information needed to write it properly is discarded at the moment of classification.

Submagic advertised 123+ languages and Choppity 97 when we read their sites on 26 August 2026. Every large caption tool lists Hindi. A reasonable person concludes that captioning Indian speech is a solved problem, uploads a Reel of ordinary Hindi-English conversation, and gets back something they cannot post.

The gap is not a coverage gap that will close as the lists grow. It is a consequence of how the interface is shaped, and a list of two hundred languages would not fix it.

What the dropdown is actually asking

A language selector performs classification. It asks you to declare which language the audio is in, then routes the audio to the model for that language. Everything downstream — the acoustic model, the vocabulary, the orthography of the output — follows from that single decision.

Code-mixed speech breaks the premise, because the question has no correct answer. A sentence that begins in Hindi, takes an English noun phrase in the middle and resolves in Hindi is not Hindi audio and is not English audio. It is both, interleaved, at a granularity finer than the unit the interface operates on.

You are therefore forced to answer wrongly, and which way you answer wrongly determines which half of your sentence gets damaged.

The two failure paths, and why neither is recoverable

Answer "Hindi" and the model transcribes to Devanagari. Your English words were spoken with English phonology, so they get written phonetically in Devanagari — "subscribe" becomes a Devanagari spelling of the sounds. If the tool then offers romanization, it transliterates that back into Roman letters and you get "sabskraib". The English word has been through two lossy conversions and arrives as a spelling no reader has ever seen.

Answer "English" and the model does what it can with the Hindi. Usually it drops it or substitutes the nearest English-sounding words, which produces a transcript that is confidently wrong rather than obviously broken — considerably harder to catch when you are skimming.

The key point is that neither is fixable downstream. The information that "subscribe" was an English word rather than a sequence of Hindi phonemes was discarded at the classification step. No amount of post-processing can recover it, because by the time post-processing runs, that distinction no longer exists in the data.

Why romanization is not the same feature

Several tools advertise romanization, and it is easy to read that as Hinglish support. It is not the same thing and the difference matters.

Romanization is a script conversion applied to a finished transcript: take Devanagari text, write it in Roman letters. It operates on characters and has no knowledge of where the words came from. Applied to a transcript that already respelled your English words phonetically, it faithfully converts that damage into Roman letters.

Writing Hinglish properly means deciding, per word, which orthography that word belongs in — English words in English spelling, Hindi words in Roman Hindi. That is a decision about word origin, which has to be made while the word is still being recognised, not afterwards.

Language selection
Asks: which language is this audio? One answer for the whole file. Cannot represent a sentence that is two languages.
Romanization
Asks: how do I write these characters in Roman script? Operates after the fact, on characters, with no memory of word origin.
Code-mix handling
Asks: for each word, which language is this and how should it be spelled? The only one of the three that can produce "subscribe" rather than "sabskraib".

The four-word test

If you are evaluating a tool, you do not need a long clip. You need one that contains the failure, and the failure is specific enough to provoke deliberately.

Record fifteen seconds in which you switch language mid-sentence at least three times, and include English words that a phonetic transliterator mangles distinctively. "Subscribe", "content", "algorithm" and "engagement" are reliable, because all four have English spellings far from their phonetic form and the damage is unmistakable.

Then read the transcript rather than watching the video. This matters more than it sounds: a well-animated caption is genuinely persuasive, and it is easy to watch a beautifully rendered clip of wrong words and come away with a good impression. The failure is in the text.

Why this is not a criticism of the big tools

It would be easy to read all this as an argument that the large caption tools are badly built. They are not — several are better-finished products than most of what else exists, and for a creator working in one language their language selector is exactly the right interface: simple, predictable, and correct for the overwhelming majority of their users.

Code-mixing is a design decision they have not made rather than a bug they have shipped. Supporting it properly means treating it as its own target from the start, which is a real cost for a market that is a small fraction of a global user base.

What is worth pushing back on is the implication that a longer language list addresses it. It does not, and a creator who reads "123+ languages including Hindi" and assumes their Hinglish Reel is covered has been misled by a number that is accurate.

Getyourvideoswatched.

Every style unlocked. No credit card.