Perspective typography
3D text captions in real perspective.
Five presets on a real 3D engine — 4×4 transforms, a perspective divide, face culling and depth order. Not a skewed layer with a drop shadow.
LumaCaption has five caption presets driven by an actual perspective engine rather than a skew: Slab, Corner and Tower write type onto the faces of a solid that turns in 3D space, and Headline and Redline print it onto a newspaper page receding from the camera. The type stays vector at every angle.
Nearly every "3D text" effect in a video tool is two-dimensional. A skew transform to suggest an angle, a stack of offset copies to suggest extrusion, a gradient to suggest a light source. It reads as 3D for about a second, and then the illusion breaks in the one place it always breaks: the perspective does not converge. Parallel edges stay parallel, and your eye knows.
These five run on a real perspective engine — 4×4 transforms, a genuine perspective divide, cuboid faces that are culled when they turn away and drawn in depth order. Type is painted onto a surface in space and projected from it, so the convergence is computed rather than imitated, and a line further from the lens is smaller because it is further away.
Three volumetric presets
Slab writes onto a tumbling block. Corner rocks between two walls of an invisible corner. Tower stacks type up a leaning tower that recedes to a vanishing point.
Two press presets
Headline and Redline print your line onto a newspaper page in perspective, its columns rushing past behind while the headline stands still.
The type stays vector
Perspective text is painted through per-strip affine fits rather than warping a bitmap, so edges stay crisp at every angle and at 4K.
Frame-for-frame identical export
The pose is a pure function of timeline position, re-derived each frame rather than accumulated — so what you previewed is exactly what renders.
The five, and what each one is
They divide into two families that share an engine and nothing else. The volumetric three put type on a solid; the press two put it on a sheet.
- Slab
- Heavy condensed caps written onto an invisible turning block, each line filling its face and tipping away over the top edge as the next lands. A hook for rap, sport and trailer cuts.
- Corner
- Two walls of an invisible corner with the block rocking between them, words switching sides and older lines greying into the depth. Architectural — for voiceover and title sequences.
- Tower
- Type stacks up a leaning tower. Each new line lands huge at the foot and shoves the rest away toward the vanishing point, while the corner swings slowly so the text winds around it.
- Headline
- The camera settles into a rushing broadsheet and your line lands as the headline, one word struck through with a highlighter.
- Redline
- Tabloid caps on a page tearing past at speed, the key word ringed in red biro. For exposés and anything that should read as evidence.
Why the box is never drawn
Slab and Corner are built on a cuboid, and the cuboid is never rendered. Only the letters exist; the form reads entirely from where the type bends across an edge.
That is a deliberate constraint rather than a missing feature. The engine can shade and occlude a solid box, and the option to turn it on exists — but these are captions, and a solid slab would black out the footage behind them. A caption that paints over the video has stopped being a caption. Leaving the object invisible keeps the shot and still reads unmistakably as one solid, because three faces of type at three different orientations can only be explained one way.
The same restraint explains an absence you will notice if you look for it: these presets carry no drop shadow and no per-glyph glow. Perspective type is painted through per-strip clips, and a shadow would be sliced at every strip boundary into visible bands. They trade the legibility shadow for the fidelity — which is why they want simple footage or a solid ground underneath, not a busy street scene.
The object does not reset between cues
This is the detail that separates the volumetric presets from an animation that merely looks 3D, and it took rebuilding to get right. The renderer runs on the whole track rather than a cue at a time.
Handled per cue, the box was wiped and started blank every couple of seconds — the viewer would watch three faces get written on, and then the object they had been reading would vanish. Real objects do not do that. Here text stays on the wall it was written on and leaves by turning out of shot, so a run of lines accumulates across several cues and older ones read sideways on a face the box turned away earlier.
Tower works the same way for the same reason, unrolled upward: nothing is ever cleared, and old lines recede toward the vanishing point and grey out rather than disappearing.
Where a perspective preset is the wrong choice
These are full-frame compositions. They are for a hook, a title sequence, a voiceover over simple footage, a faceless channel, a trailer beat — anywhere the type is allowed to be the picture.
They are the wrong choice over a talking head, because the composition covers the frame you are asking people to watch. They are also the wrong choice for dense, fast narration: type on a turning surface takes marginally longer to read than type sitting flat, and that cost compounds across a three-minute clip.
The usual build is to use one for the first cue and hand off to a flat style for the rest — any single line can wear its own look, so that is one click rather than a re-edit.
How it works
Bring audio or video
A talking clip, a voiceover, or a recording made straight in. Word timings come from the audio. The video file is never uploaded.
Pick one of the five
Slab, Corner or Tower for volumetric type; Headline or Redline for the press look. Each previews on your own footage.
Export at 4K if it matters
The type is vector at every angle, so a 4K export is genuinely sharp. 1080p is free; 4K is on the paid plans.
Common questions
Is this real 3D or a skew effect?
Real 3D. There is a perspective engine underneath — 4×4 transforms, a genuine perspective divide, cuboid faces culled when they turn away and drawn in depth order. That is why parallel edges converge and why a line further from the lens is genuinely smaller. A skew cannot produce either, which is the specific tell that gives fake 3D away.
Which presets use it?
Five. Slab, Corner and Tower write type onto a solid turning in space; Headline and Redline print it onto a newspaper page in perspective. The first three sit in the Big Type lane and the press pair in Paper & Print, but all five share the same engine.
Does the text get blurry at an angle?
No. Perspective type is painted through per-strip affine fits rather than by warping a rendered bitmap, so it stays vector at every angle — which is also why a 4K export of one of these is genuinely 4K rather than an upscale.
Why do these presets have no drop shadow?
Because perspective text is painted through per-strip clips, and a shadow would be sliced at every strip boundary into visible bands. They trade the legibility shadow for fidelity, which means they want simple footage or a solid ground behind them rather than a busy shot.
Can I use one for a whole video?
You can, but they are built as full-frame compositions and they are harder to read at length than flat type. The build that works is one of these on the hook and a flat style for everything after — any single line can carry its own look, so switching costs one click.
