AI-Assisted Subtitling and Timing
The cooking show was forty-four minutes long, the deadline was the next morning, and the engine had already drafted every subtitle. A media localizer opened the Spanish file feeling almost lucky: the translation read beautifully, every line faithful, every joke intact, the chef's fast patter rendered word for word. She watched the first three minutes to spot-check and felt her stomach drop. The chef would say a single quick sentence and three full lines of perfect Spanish would flash up and vanish before she could finish reading the first one. A faithful translation she could not read. By minute two a subtitle was bleeding across a hard cut to a wide shot, lingering over the next speaker's mouth. By minute four two cues overlapped on screen at once, stacked four lines deep, covering half the food. Nothing was mistranslated. Everything was unwatchable. The engine had done exactly what engines do: it had protected the meaning and ignored the clock, the box, and the eye. This lesson is about that clock, that box, and that eye, the structural laws of the subtitle that have nothing to do with whether the words are right and everything to do with whether a human being can actually read them in the seconds they are on screen.
The Subtitle Is Not a Translation
Start with the definition, because the whole lesson hangs on it. A subtitle is not a translation of a line of dialogue. It is a line of translated text displayed on screen for a measured interval, timed to the speech it represents, fitted into a box of fixed width, and constrained by how fast a human being can actually read. The translation is one of its properties. It also has an in-time and an out-time, a duration measured to the frame, a character budget per line, a limit on how many lines may share the screen, and a relationship to the cut of the picture. A subtitle with perfect words and wrong timing is not a slightly worse subtitle. It is a broken one, because the words the viewer cannot finish reading are no better than words that were never written.
This is the same shape you met one lesson earlier with placeholders and tags, and it is worth naming the parallel out loud. There, a software string was language carrying structural cargo, and a fluent machine-translation (MT) engine, any system that converts source text to target text without a human writing the words, would rewrite or destroy the cargo because the cargo was invisible to a process built only to make sentences sound good. A large language model (LLM), a general-purpose text predictor that translates as a side effect of predicting plausible next words, does the same. The subtitle is that problem with a new and harsher kind of cargo: not just placeholders and tags, but time itself. A button has a pixel budget. A subtitle has a pixel budget and a second budget, and the second budget is the one the engine is structurally blind to.
The engine protects the meaning of a subtitle and ignores the clock, the box, and the reading eye. A perfectly translated subtitle the viewer cannot read in time is, to that viewer, simply broken.
Roughly 70% of language-service providers (LSPs) now offer subtitling, and a large and growing share of it is machine-first: the engine drafts every cue before a human opens the file, exactly as it now pre-populates every segment in the translation-management system (TMS). That means the engine that mistranslates a placeholder is now also being handed the timecodes, the line breaks, and the reading rate, and it understands none of them. The opportunity is real, the speed is real, and the failure modes are specific, repeatable, and almost entirely invisible to a linguist who reads only for meaning. Learning subtitling in an MT-first world is learning to read for the things the words being right will never tell you.
Reading Rate: The First Law
The governing constraint of all subtitling, the one from which most of the others follow, is the reading rate: the speed at which a viewer can comfortably read subtitle text while still watching the picture, measured in characters per second, abbreviated CPS. Reading rate is not a style preference and it is not negotiable by the translator's taste. It is a measured property of human reading, and a subtitle that exceeds it forces the viewer into a bad choice: stop watching the image to finish reading the text, or give up on the end of the line and miss what it said. Either way the subtitle has failed at the one job it exists to do.
The widely used professional norm sits around 15 to 17 characters per second for general adult content, and lower, often nearer 12 CPS or below, for children's programming, where younger viewers read more slowly. CPS is computed simply: take the number of characters in the cue, divide by the cue's duration in seconds. A two-line subtitle of 74 characters that is on screen for 4 seconds runs at roughly 18.5 CPS, which is already over the comfortable line for adult content. The same 74 characters given 5 seconds run at about 14.8 CPS, comfortably readable. Nothing about the words changed. The only thing that changed was the time they were given, and that single number decided whether the subtitle was watchable. This is the first habit of the AI-assisted subtitler: you do not read the subtitle, you read its CPS, because the prose being good tells you nothing about whether it is on screen long enough to be read.
Why the Engine Blows the Reading Rate
An MT or LLM engine left to itself does the precise opposite of what reading rate requires, and it does so for a reason baked into what the engine is. The engine optimizes for completeness and fluency: a full, natural, faithful rendering of everything that was said. That is the right instinct for a document and the wrong instinct for a cue with 1.8 seconds of screen time, because the most faithful rendering is frequently the longest one, and length is exactly the thing the clock will not forgive. The engine does not know there is a clock. It is handed the source line "Well, I mean, honestly, at the end of the day, what I'm really trying to say is that we should probably just go," and it produces a complete and fluent target rendering of all twenty-two words, because dropping any of them would feel like an inaccuracy to a process trained to be accurate. The cue is on screen for two seconds. The result is a perfect translation racing past at 30-plus CPS, which is to say an unwatchable one.
Here is the trap that makes this dangerous rather than merely annoying: the overrun is invisible to a reviewer reading the file as text. Open the subtitle in an editor and read it, and every line is fluent, faithful, well-formed Spanish or German or Japanese. There is nothing to flag, because the error is not in the words. The error is in the relationship between the character count and the duration, and that relationship is only visible if you compute it or if you watch the subtitle play at speed. A linguist proofreading the transcript would approve every line. The file would still be unwatchable. This is the subtitling version of the program's spine: fluent is not correct, and here "correct" includes a property the prose cannot reveal.
Condensation: The Craft the Engine Skips
The professional answer to a reading-rate overrun is condensation: conveying the meaning of the line in fewer characters so it fits the available time at a readable rate. Condensation is not summarizing and it is not dumbing down. It is the disciplined removal of what the cue can afford to lose, filler words, redundant hedges, repetition, an aside the picture already conveys, while protecting what carries the meaning, the tone, and the intent. That fast cooking-show line becomes, in a good subtitle, the equivalent of "Honestly, we should just go." Same meaning, same register, a third of the characters, readable in the time the cue allows. The viewer never feels the cut, because nothing they needed was cut.
Condensation sits at the exact collision of two jobs. It is a linguistic decision, what to keep and what to sacrifice so the meaning survives, which is the linguist's judgment and no one else's. And it is bounded by a hard numeric constraint, the characters that fit in the seconds available, which is arithmetic the tooling can compute and enforce. The AI-assisted subtitler's leverage is to let the engine and the tooling do the parts they are good at, draft a full rendering and compute the CPS, and then do the part only a human can: decide which words to lose so the line reads in time without losing what the scene needs. An engine asked to "translate this subtitle" will not condense, because condensing looks like inaccuracy to it. An engine told "this cue is 1.8 seconds and 42 characters per line, produce a faithful rendering that fits at 16 CPS" can produce a useful shorter draft, but the call on whether the shortened version still says what the chef meant belongs to the person who understands the chef, the scene, and the language. The number is the engineer's; the judgment is the linguist's; and the engine respects only the constraint it is told about.
Condensation is a linguistic decision inside an arithmetic box. The engine will draft a complete translation and overrun the clock; only a human decides which words the cue can afford to lose.
Line Length and Line Count
Reading rate governs time. The second family of laws governs space: how wide a line may be, how many lines may share the screen, and where a line is allowed to break. These are conventions, not physics, but they are stable enough across the industry to function as rules, and the engine respects none of them because it produces a subtitle as a single undifferentiated run of prose.
The first rule is line length: a single subtitle line is conventionally limited to a maximum number of characters, commonly around 37 to 42 for Latin scripts, with 42 the most familiar ceiling. The limit exists because a line wider than that either does not fit the safe area of the frame or forces the eye to travel too far horizontally to track comfortably. The second rule is line count: no more than two lines on screen at once. Two lines is the ceiling because a third line crowds the picture, pushes subtitles up into the action, and gives the eye more text than it can absorb in the cue's duration. A subtitle that needs more than two lines of text at a readable rate is a subtitle whose source line is too long for its time, which sends you back to condensation or to splitting the dialogue into more cues.
An engine handed a long source line and told to translate it produces one long target string with no concept of a 42-character ceiling or a two-line maximum. The tooling wraps that string to fit the box, and where it wraps is wherever the text happens to run out of width, which brings us to the rule the engine breaks most quietly of all.
Where a Line Breaks: Segmentation Inside the Cue
When a subtitle spans two lines, the place it breaks is a linguistic decision, not a mechanical one, and getting it wrong degrades readability even when every word is correct. A good line break, the kind a professional subtitler makes on purpose, falls at a natural grammatical boundary: after a complete clause or phrase, so each line can be parsed as a self-contained unit while the eye moves to the next. A bad line break splits a grammatical unit that the reader then has to hold incomplete until the second line resolves it. Breaking an article from its noun ("the" at the end of line one, "agreement" at the start of line two), a preposition from its object, an auxiliary from its verb, or a name across two lines forces the reader to suspend comprehension across the break, which costs reading time the cue may not have. The viewer feels it as friction without knowing why.
This is segmentation at the line level: the act of dividing a cue's text into well-formed lines that break at sense boundaries. An engine producing fluent prose has no model of the two-line box and no notion of a grammatical break point, so when its output is wrapped to fit, the break lands wherever the 42nd character happens to fall, which is frequently mid-phrase. The defense is the same locked-structure discipline as everywhere else in the pipeline: the line breaks are decided by a human reading for sense, or by tooling that breaks at clause boundaries, never left to fall where the raw text width dictates. A subtitle can be perfectly translated, perfectly timed, and still read badly because it broke "the new safety" / "regulations apply" instead of "the new safety regulations" / "apply to all staff." The words are right. The shape is wrong, and the shape is part of the job.
Segmentation Across Shots: The Cut Is a Character
There is a second, larger sense of segmentation that matters even more than line breaks, and it is the one the engine is most completely blind to. This is segmentation at the cue level: deciding where one subtitle ends and the next begins across the flow of speech and, crucially, across the cuts of the picture. The dialogue is a continuous stream; the subtitle file is a sequence of discrete cues; and the work of deciding where to slice the stream into cues is structural, visual, and rhythmic work that has almost nothing to do with the words and everything to do with the edit of the film.
The cardinal rule of cue segmentation is that subtitles should respect shot changes, not straddle them. A shot change, a hard cut from one camera angle or scene to another, is a powerful visual event; the eye instinctively re-reads any text on screen when the image changes, because it assumes new image means new information. A subtitle that begins before a cut and continues after it forces the viewer to re-read it at the cut, wasting reading time and producing a subliminal sense of wrongness. A subtitle that ends exactly on the cut, or starts cleanly after it, rides the edit instead of fighting it. Professional subtitling aligns cue boundaries to shot boundaries wherever the dialogue allows, ending a subtitle a frame or two before a cut and starting the next one cleanly on or after the new shot.
An engine does not see the picture. It sees a transcript, or it sees a source subtitle file with cues already segmented by whoever made the original-language version, and it translates cue by cue. If the original cues were well segmented and the engine touches only the text, the segmentation survives by inheritance. But the moment the target language's different length or word order would be better served by re-segmenting, merging two short cues, splitting one long one, shifting a boundary to respect a cut the engine cannot see, the engine has no basis to do it, because the cut exists in the video and the engine has only the text. Worse, an engine allowed to reflow the file rather than edit cue text in place can shift or merge cues blindly, dragging boundaries off the shot changes the original respected. Cue segmentation is video work. It belongs to the subtitler watching the picture, and it is one of the clearest reasons a fluent engine cannot finish a subtitle file on its own.
Sync: The Second Clock
Reading rate is one clock, the duration a cue needs to be readable. Sync is the other: the alignment of each subtitle's appearance and disappearance with the speech and the picture. Sync is the property that most loudly separates professional subtitling from amateur work, because human viewers are exquisitely sensitive to the timing of text against speech even when they cannot read the language. A subtitle that appears a beat before the line is spoken spoils a reveal or a punchline by announcing it early. A subtitle that lingers after the speaker has stopped, or hangs over a new speaker's mouth, reads as sloppy. A subtitle that bleeds across a hard cut, as in the cooking show, jars every time.
Sync is governed by a set of professional norms that function as hard constraints. There is a minimum duration a cue must stay on screen even if the line is very short, so a one-word subtitle does not flash and vanish before the eye can land on it, commonly something on the order of a second or a little less. There is a maximum duration beyond which a static subtitle starts to feel stuck and the eye re-reads it pointlessly, often around six seconds. There is a minimum gap between consecutive cues, a couple of frames of blank, so the viewer's eye registers that one subtitle ended and another began rather than seeing a continuous smear of changing text. And there are the lead-in and lead-out tolerances: how close to the actual start of speech a subtitle should appear, and how long after it should clear. These numbers are not the linguist's invention; they are media-localization conventions, and they live in the file as data, not as prose.
The Timecodes Are Locked Structure
Every one of those timings is encoded in the subtitle file itself, in formats like SRT (the simple SubRip format) and WebVTT (the web standard), as an in-timecode and an out-timecode attached to each cue. A cue in an SRT file is a number, a line reading 00:01:12,480 --> 00:01:15,920, and then the one or two lines of subtitle text. The timecodes are the sync. They are machinery, exactly like a placeholder or a tag, and they obey exactly the same rule: when an engine translates a subtitle file, it must touch only the text inside each cue and leave the timecodes and the cue structure untouched. The text is language. The timecodes are an instruction to the player about when to show that language, and they must arrive in the target file byte-for-byte as they were in the source, or the carefully built sync of the original is destroyed.
The failure mode here is familiar from the placeholder lesson and just as silent. An engine handed the whole file as text, rather than told to edit only the cue text in place, can do several fluent and fatal things. It can "tidy" a timecode it mistook for malformed text. It can drop the blank lines that separate cues, merging two cues into one and collapsing their timings. It can renumber or reorder cues. It can translate the format's own keywords or reflow the cue boundaries to suit the target prose, shifting every downstream timecode out of sync. And as always, the result reads perfectly as text: open it and every subtitle line is good. Play it against the video and the words land a second early, or two cues are stacked on screen, or the timing drifts further out with every minute. The words protect the meaning; the timecodes protect the sync; and only the words are something a prose engine understands. The discipline is to lock the timecodes and the cue structure so the engine and the post-editor can change the text and physically cannot change the timing, and to validate after the fact that the set of cues and their timecodes in the target match the source exactly.
Timecodes are placeholders for time. The engine that translates the costume of a word will reflow a cue and break the sync, and the meaning-reader never sees the wound because the words are still perfect.
Where AI Genuinely Helps, and Where It Breaks
It would be a mistake to read all of this as an argument against using the engine. The engine is genuinely useful in subtitling, and the AI-assisted subtitler's skill is knowing exactly which parts of the job to hand it and which parts to keep. Drawing that line cleanly is the difference between speed that compounds and speed that ships unwatchable files.
The engine helps most with the raw linguistic transfer inside a cue whose structure is already sound. Given a well-segmented source subtitle file with good cue boundaries and correct timecodes, an MT or LLM engine can draft a fluent target rendering of each cue's text fast, and if the tooling tells it the per-line character limit and the cue duration, it can even produce a draft that is aware of the budget and attempts to fit it. It can suggest a shorter phrasing when asked to condense to a target CPS. It can carry a glossary so recurring terms and character names stay consistent across a forty-four-minute file. It can take the first, complete pass that the human then shapes, which is the same MTPE (machine-translation post-editing) economics the rest of the program runs on: the engine lifts throughput, the human owns the quality.
The engine breaks at every point where the job depends on something outside the text. It cannot judge condensation, because it does not know which words the scene can afford to lose. It cannot segment cues against the picture, because it cannot see the cuts. It cannot fix sync, because timing is a property of the video and the file, not of the prose. It cannot reliably honor a reading rate it is not explicitly given, and even when given one, its sense of "good enough to read" is not the viewer's. And it will, unless physically prevented, reflow the very timecodes and cue structure that hold the whole thing together. The line is clean: the engine drafts the language inside a cue; the human owns the time, the box, the cut, and the eye. Hand it the language and it accelerates you. Hand it the structure and it ships you a beautiful, unwatchable file.
A Worked Subtitle Post-Edit
Theory settles when you watch it applied to a single cue. Here is one, drawn from the kind of fast dialogue that breaks engines. The source is English; the target is Spanish; the cue's timecodes give it 2.0 seconds on screen, and the line limit is 42 characters with a two-line maximum and a reading-rate target of 16 CPS. The speaker is talking quickly.
Source cue (2.0 seconds on screen):
00:04:18,200 --> 00:04:20,200- "Look, I really don't think we should sign this agreement until the lawyers have actually had a chance to read it."
The engine's draft: a complete, fluent, faithful Spanish rendering of every word, returned as a single long run of text, something equivalent to "Mira, realmente no creo que debamos firmar este acuerdo hasta que los abogados hayan tenido la oportunidad de leerlo." Count it: that is well over 100 characters. Forced into the cue, the tooling wraps it onto three or four lines, blowing the two-line maximum, and at 2.0 seconds of screen time it runs north of 50 CPS, more than triple the readable rate. As text in an editor it is flawless: every word is right, the register is right, nothing is mistranslated. As a subtitle it is unreadable, oversized, and would flash three or four lines of text past the viewer in two seconds. The engine did its job, protecting the meaning, and failed the subtitle on every structural axis at once.
The post-edit, step by step. The first move is not to read the words; it is to read the constraint. The cue is 2.0 seconds, so at 16 CPS the readable ceiling is about 32 characters total, across at most two lines. The full rendering is over 100. This is not a polish job; it is a condensation job, and the budget is brutal: roughly two-thirds of the characters have to go while the meaning survives. So the post-editor asks what the line actually needs to convey. The intent is a single, clear objection: do not sign until the lawyers have read it. The hedges ("Look," "I really don't think," "actually had a chance to") are filler the speaker used for rhythm, not content; the picture and the performance already carry the speaker's reluctance. The core is "Don't sign until the lawyers read it."
The condensed Spanish becomes something equivalent to "No firmemos hasta que los abogados lo lean." That is around 43 characters. Still slightly over the two-second budget for one line, so the post-editor either splits it cleanly across two lines at a grammatical boundary, "No firmemos / hasta que los abogados lo lean," breaking after the verb so each line is a parseable unit, or, if even that runs hot against the clock, trims once more to "No firmemos sin que lo lean los abogados," tightening the structure. The reading rate now sits in the readable range, the line obeys the 42-character limit, the break (if used) falls at a sense boundary, and the meaning, the objection and its reason, is fully intact. Nothing the scene needed was lost. Everything the clock demanded was honored.
What the post-editor verified that reading alone would never reveal. Notice everything in that edit that had nothing to do with whether the Spanish was "good." The CPS was computed against the cue duration. The character count was checked against the line limit. The line count was kept at two. The break point was placed at a grammatical boundary, not where the width ran out. And, the step that frames the whole file, the timecode 00:04:18,200 --> 00:04:20,200 was never touched: the cue's sync to the speaker's mouth and to the surrounding shots was preserved exactly, because the post-editor edited the text inside the locked cue and left the timing alone. The engine's draft was the starting point that saved the typing. The post-edit was the judgment, linguistic and structural at once, that turned a flawless translation into a watchable subtitle. That is the division of labor the rest of this lesson described, executed on one cue, and it is the move you repeat several hundred times across a forty-four-minute file.
The Quality Check Before Delivery
One cue done well is not a delivered file. Before a subtitle file ships, the AI-assisted subtitler runs a structural pass that no amount of reading the prose can substitute for, because every check in it tests a property the words cannot reveal. Concretely, the pass confirms that no cue exceeds the reading-rate target, flagging every cue whose CPS runs hot for re-condensation; that no line exceeds the character limit and no cue exceeds two lines; that every two-line cue breaks at a sense boundary; that cue durations sit between the minimum and maximum and the gaps between cues are present; that cue boundaries respect shot changes wherever the dialogue allowed; and, the locked-structure check, that the count, order, and timecodes of the target cues match the source exactly, proving the engine and the post-edit changed only the text and never the timing. Many of these are arithmetic the tooling computes and flags automatically, which is the right division: the tool catches the measurable violations, and the human spends attention on the judgment calls, condensation and segmentation, that the tool can flag but cannot decide. A file that passes that structural pass and reads well is a file you can deliver. A file that only reads well is a guess.
Key Takeaways
- A subtitle is not a translation; it is timed, boxed, reading-rate-bounded text where the words are only one property. An MT or LLM engine protects the meaning and is blind to the clock, the box, the cut, and the reading eye, so a perfectly translated subtitle can be completely unwatchable, and that failure is invisible to anyone reading the file only as text.
- Reading rate, measured in characters per second (CPS), is the first law: roughly 15 to 17 CPS for adult content and lower for children's. Compute it as characters divided by cue duration. The same words are readable or unreadable depending only on the seconds they are given, so you read a cue's CPS, not just its prose.
- Engines overrun the reading rate because they optimize for complete, fluent renderings, and the most faithful rendering is often the longest. The fix is condensation: a linguistic decision about which words a cue can afford to lose, bounded by the arithmetic of characters per second. The engine drafts and computes the budget; only a human decides the cut.
- Line length (commonly up to about 42 characters) and a two-line maximum govern the subtitle's space. Where a two-line cue breaks is a linguistic decision: good breaks fall at grammatical boundaries so each line parses as a unit, bad breaks split a phrase and cost reading time. Engines produce one undifferentiated run and break wherever the width runs out.
- Cue-level segmentation must respect shot changes, never straddle them, because the eye re-reads text at a cut. This is video work the engine cannot do because it sees only text, not the picture. An engine allowed to reflow a file rather than edit cue text in place will drag cue boundaries off the shot changes the original respected.
- Sync is the alignment of each cue's appearance and disappearance with speech and picture, governed by minimum and maximum durations, minimum gaps between cues, and lead-in/lead-out tolerances. These timings live as in- and out-timecodes in formats like SRT and WebVTT, and they are locked structure, exactly like a placeholder.
- When an engine translates a subtitle file it must touch only the cue text and leave the timecodes and cue structure byte-for-byte intact; an engine allowed to reflow the file can tidy a timecode, merge cues, or shift boundaries and silently destroy the sync while every line still reads perfectly. Lock the timing, edit the text, then validate that target cues and timecodes match the source.
- The division of labor is clean: hand the engine the language inside a cue and it accelerates you with an MTPE-style first draft; keep the time, the box, the cut, and the eye for the human. Deliver only after a structural pass that checks CPS, line length and count, break points, cue durations and gaps, shot-change alignment, and exact timecode match, because every one of those tests a property the prose cannot reveal.
Skill.re