Auto-reframe
Auto-reframe is an AI feature that converts a video from one aspect ratio to another by tracking the subject (typically the active speaker's face) across frames and re-cropping the output to keep the subject centered. The dominant use case is converting 16:9 source footage into 9:16 vertical clips for TikTok, Reels, and YouTube Shorts.
Without auto-reframe, converting a 16:9 source to 9:16 either crops to the center third (losing the subject if they move out of center) or letterboxes (wasting 40% of the visible canvas). Auto-reframe solves this by detecting the active speaker per frame and adjusting the crop window dynamically. The result feels natively 9:16 rather than mechanically cropped.
Modern auto-reframe handles multi-speaker scenes (two-host podcasts, interviews) by switching the crop window between speakers based on audio activity. Quality varies between tools — some auto-reframe products produce visible jitter when the speaker moves; others smooth the crop trajectory. Clipzing.com's auto-reframe targets sub-frame jitter and handles multi-speaker switching cleanly.
Auto-reframe is now table-stakes for any AI clip generator. The differentiator in 2026 is quality of the crop trajectory (smooth vs jittery) and accuracy of multi-speaker switching, not whether the feature exists.
Related terms
Aspect ratio
Aspect ratio is the proportional relationship between a video's width and height, expressed as width:height. The dominant short-form ratio in 2026 is 9:16 (vertical, 1080×1920) used by TikTok, Instagram Reels, and YouTube Shorts. Long-form video uses 16:9 (1920×1080). Square 1:1 is used by LinkedIn and Instagram feed.
Short-form video
Short-form video is the category of video content typically 15 to 90 seconds long, formatted in 9:16 vertical aspect ratio, designed for thumb-scrollable mobile feeds. The dominant short-form platforms in 2026 are TikTok, Instagram Reels, YouTube Shorts, and Snapchat Spotlight. Long-form video (3+ minutes) lives on YouTube and traditional video platforms.
B-roll
B-roll is supplementary footage cut into a video to support or illustrate the main A-roll (the speaker on camera). In short-form video, B-roll is typically 1–4 second inserts of relevant imagery, gameplay, screen recordings, or stock footage that visualize what the speaker is describing. Used well, B-roll boosts retention; used to cover the punchline, it kills replay.
Transcription
Transcription is the process of converting spoken audio into written text. In AI clip generators, transcription is the foundation step — every other feature (auto-captions, viral-clip scoring, semantic search across the source) depends on accurate transcription. Modern speech-to-text models hit 95%+ accuracy on clear English audio with native speakers.