Key Facts
- A clear still frame does not prove a caption stays clear during motion.
- Check labels and spoken captions together.
- Use the actual words, including the longest line.
- Record collisions by shot and timestamp.
Which Text Deserves Space In The Frame?
Spoken captions let viewers read the dialogue, while object labels identify details in the picture. Add a heading only when it introduces something the other text does not. If several elements repeat the same phrase, remove the redundant text before making everything smaller.
In an author-created example, a fictional Short shows how to fold a fabric napkin into a pocket. The spoken instruction is “Tuck the lower corner underneath.” A label pointing to the lower corner may help briefly. A large heading saying “TUCK IT” adds little while occupying the same useful space.
Plan where the hands travel. The bottom of the frame might look empty in the opening shot but become the most important area once folding starts.
Build the framing plan first: https://dreamwild.ai/guides/vertical-video-composition/
How Do You Run A Useful Collision Test?
Use the actual exported draft with captions included. Play it on a phone at a comfortable viewing distance, then inspect moments where the longest lines appear. A desktop preview stretched across a monitor can hide a size problem.
The checklist below is an author-created test record for the napkin example. Its retest results are fictional and do not establish readability on other devices.
Bring the exported draft and find its longest caption. For each device and player you test, note any overlap with the action or controls, then replay the moment after correcting it.
Scroll this table sideways to see every column.
| Review Moment | Collision Found | Revision | Retest |
|---|---|---|---|
| Fold begins | Caption covers the corner being lifted | Move the reading area above the hands | Corner remains visible throughout lift |
| Longest instruction | Last word touches the edge after line wrapping | Split at a natural phrase boundary | Both phrases readable without pausing |
| Object label appears | Arrow crosses the spoken caption | Show the label before the next caption phrase | Reading order is clear |
| Result held | Viewing controls overlap the finishing label | Move the label beside the pocket | Label clear on the tested surface |
| Final playback | Caption jumps between three positions | Keep one position for the folding sequence | Eye movement feels predictable |
Record the device, viewing surface and video version beside your results. “Checked on phone” is too vague when a later crop or caption edit changes the file.
What Should You Change First When Text Does Not Fit?
Shorten the wording without changing the instruction. “Fold the bottom corner up” may communicate the same action more directly than “You now want to take the corner at the bottom and fold it upwards.”
Next, adjust phrase breaks. Keep words that belong together on the same line when possible. If the caption changes halfway through a term, the reader must reconstruct it while watching the hands.
Then try position and background treatment. A small solid backing can make text readable across a moving scene, but it also covers part of the picture. Give it deliberate space. Reducing the font size should not be the automatic first response.
Turn the script into caption phrases before detailed placement: https://dreamwild.ai/guides/spoken-words-to-captions/
How Do You Test More Than One Viewing Situation?
Start with the intended player and test a contrasting screen size if available. Watch with the interface visible as well as during unobstructed playback. Use ordinary brightness and avoid holding the screen unusually close.
You do not need to invent a claim about every phone. Write down the tests you actually performed. If a second device is unavailable, note that limit in your production record and use a conservative layout with fewer edge-dependent details.
Check the video with sound muted. Captions may read well individually yet move too quickly to follow the instruction. Then check with sound to find timing mismatches.
When Is The Placement Ready?
Approve the text when a viewer can read it, identify its target and watch the necessary action in the same pass. If a dense diagram needs sustained attention, let the caption finish before asking the viewer to inspect the diagram.
Save the approved caption position for similar scenes, then recheck it when the subject, crop or interface changes. In the next folding demonstration, the hands may occupy a different part of the frame.
Frequently Asked Questions
Is There One Safe Margin For Every Phone?
This guide does not prescribe one. Review the actual video on the viewing surfaces you intend to use and leave room for their visible controls.
Should Captions Follow A Moving Person?
Only when that movement improves comprehension. A stable caption area is often easier to review and less likely to collide with the action.
Can I Put Labels And Captions On Screen Together?
Yes, when both remain readable and serve different purposes. Shorten or stagger them if the viewer must choose which one to read.
Do I Need To Test The Export Again?
Yes. The exported file is the version the viewer receives, so check its text size, timing and crop even if the editor preview looked correct.
Your next step
Put the idea to work.
Explore DreamWild's studio for reviewing captions alongside your scenes
Explore the creative studio https://dreamwild.ai/creative-studio/