Writing

What we measured while building this.

Notes from building a captioning engine — the standards we read, the things we got wrong, and the numbers behind decisions that are usually described as taste.

What gets written here, and what does not

Captioning looks like a solved problem from the outside. Transcribe the audio, put the words on the screen, done. Almost every hard part of it is invisible until you have shipped something and watched it fail on real footage — type that dissolves against a bright sky, a script whose vowel marks get clipped because the line box was measured for Latin, speech that switches language halfway through a sentence and has nowhere to go in an interface built around a dropdown.

Those are the things worth writing down, so that is what these posts are. Each one carries a number, a threshold or a specification we actually use, because a post that only describes a problem is an essay and a post that quotes the constant is something you can check, argue with, or borrow.

What you will not find here is a product announcement dressed as an article, a listicle of tips, or a claim we cannot show the workings for. Where we got something wrong — and the bright-footage post is mostly an account of getting it wrong for a while — the wrong version is described too, because the failed approach is usually the more useful half.

Getyourvideoswatched.

Every style unlocked. No credit card.