One source, three caption treatments

A first-party inspection of the public Clyps demo files: the same NASA interview excerpt rendered with Minimal, Bold, and Karaoke captions. Measurements checked September 7, 2026.

Last updated: September 7, 2026

What this study measures

This is a file-level inspection of one short sample, not a benchmark of clipping accuracy, processing speed, audience engagement, or competing products. The source is an existing interview and the three outputs use the same excerpt. Caption styling is the variable visitors can compare.

The sample is useful for judging where captions sit, how individual words appear over time, and how a landscape interview reads inside a vertical frame. It cannot establish the quality of every speaker, language, recording, or scene the product may encounter.

Source and method

The source is NASA Astronaut Don Pettit Talks with AstroKobi, published January 8, 2025. We used the passage from 04:02.520 to 04:09.200. The underlying NASA footage is public domain in the United States; NASA and the speakers do not endorse Clyps.

The published sample outputs are generated by the Clyps landing-media rendering script from a fixed source excerpt and caption timing. We inspected the shipped MP4 files with ffprobe, recording their video dimensions, frame rate, duration, codec, and byte size. The downloadable JSON below contains the measured values. We did not round byte counts or substitute estimated values.

Measured files

All four files contain H.264 video and AAC audio. Output durations are 6.68 seconds. The source container duration is 6.690017 seconds. The small difference in container duration is not evidence of a different selected spoken passage.

Published demo files inspected September 7, 2026
FileDimensionsFrame rateBytes
Source1280 × 72060000/1001 fps1,269,431
Minimal540 × 96030 fps886,886
Bold540 × 96030 fps920,134
Karaoke540 × 96030 fps921,930

How to read the comparison

Play each treatment through the end before choosing a style. Minimal keeps the text quieter; Bold gives the caption more visual weight; Karaoke changes the emphasis as the words progress. Inspect whether the treatment helps you follow the sentence while preserving the speaker’s face and expression.

The output files use 540 × 960 to keep this public website demonstration small. That is the resolution of these demonstration assets, not a statement that customer exports are limited to 540p. Clyps supports 720p, 1080p, and 4K export choices as described on the product pages.

Reproduce the inspection

Download any of the MP4 files linked below and run ffprobe -v error -show_entries stream=codec_name,width,height,r_frame_rate:format=duration,size -of json FILE.mp4. Replace FILE.mp4 with the downloaded filename. The dimensions, rates, duration, and byte size should match the table and raw JSON for these files.

To evaluate the product for your own work, repeat the editorial review with one of your recordings. Check names, pauses, crop changes, and whether the selected passage stands on its own. This sample provides observable output; it does not replace that review.