One source, three caption treatments
A first-party inspection of the public Clyps demo files: the same NASA interview excerpt rendered with Minimal, Bold, and Karaoke captions. Measurements checked September 7, 2026.
Last updated: September 7, 2026
What this study measures
This is a file-level inspection of one short sample, not a benchmark of clipping accuracy, processing speed, audience engagement, or competing products. The source is an existing interview and the three outputs use the same excerpt. Caption styling is the variable visitors can compare.
The sample is useful for judging where captions sit, how individual words appear over time, and how a landscape interview reads inside a vertical frame. It cannot establish the quality of every speaker, language, recording, or scene the product may encounter.
Source and method
The source is NASA Astronaut Don Pettit Talks with AstroKobi, published January 8, 2025. We used the passage from 04:02.520 to 04:09.200. The underlying NASA footage is public domain in the United States; NASA and the speakers do not endorse Clyps.
The published sample outputs are generated by the Clyps landing-media rendering script from a fixed source excerpt and caption timing. We inspected the shipped MP4 files with ffprobe, recording their video dimensions, frame rate, duration, codec, and byte size. The downloadable JSON below contains the measured values. We did not round byte counts or substitute estimated values.
Measured files
All four files contain H.264 video and AAC audio. Output durations are 6.68 seconds. The source container duration is 6.690017 seconds. The small difference in container duration is not evidence of a different selected spoken passage.
| File | Dimensions | Frame rate | Bytes |
|---|---|---|---|
| Source | 1280 × 720 | 60000/1001 fps | 1,269,431 |
| Minimal | 540 × 960 | 30 fps | 886,886 |
| Bold | 540 × 960 | 30 fps | 920,134 |
| Karaoke | 540 × 960 | 30 fps | 921,930 |
How to read the comparison
Play each treatment through the end before choosing a style. Minimal keeps the text quieter; Bold gives the caption more visual weight; Karaoke changes the emphasis as the words progress. Inspect whether the treatment helps you follow the sentence while preserving the speaker’s face and expression.
The output files use 540 × 960 to keep this public website demonstration small. That is the resolution of these demonstration assets, not a statement that customer exports are limited to 540p. Clyps supports 720p, 1080p, and 4K export choices as described on the product pages.
Reproduce the inspection
Download any of the MP4 files linked below and run ffprobe -v error -show_entries stream=codec_name,width,height,r_frame_rate:format=duration,size -of json FILE.mp4. Replace FILE.mp4 with the downloaded filename. The dimensions, rates, duration, and byte size should match the table and raw JSON for these files.
To evaluate the product for your own work, repeat the editorial review with one of your recordings. Check names, pauses, crop changes, and whether the selected passage stands on its own. This sample provides observable output; it does not replace that review.