ClipSpark AI

Blog

Choosing the shape of a clip: 9:16, 1:1 or 16:9

Vertical is not automatically right. What a 16:9 two-shot loses when it is cropped to 9:16, when a square export is the better compromise, and when to keep the frame you recorded.

Vertical is the default answer, and for a lot of clips it is the right one. It is also the answer that quietly destroys the most footage, because a 9:16 export of a 16:9 recording is not a resize. It is a crop, and a crop throws pixels away permanently.

How much a vertical crop actually removes

The arithmetic is worth doing once. Take a 16:9 frame and crop it to 9:16 without losing any height. The result is just under a third of the original width — about 32% of it. Two thirds of what the camera recorded is gone, and it is gone from the sides, which is exactly where a second person sits.

That is the whole problem with clipping a conversation. A two-shot podcast recorded 16:9 has one person on the left and one on the right, and a centred vertical crop keeps the gap between them: half of each face, neither of them properly, both of them talking to somebody off-screen.

Following the speaker

ClipSpark has an optional setting for this. Instead of holding one crop for the whole clip, the vertical frame moves onto whoever is speaking and switches when the other person takes over. On a two-person show it is usually the difference between a usable clip and an unusable one.

It is worth being precise about how it decides, because that determines when it works. It watches mouth movement on camera. It is not analysing the audio and matching a voice to a person; it is looking at faces and asking which mouth is moving. Three consequences follow, and none of them are hidden:

  • It needs faces it can see, roughly front-on. A guest turned away from the camera, sunglasses, a hand over the mouth, or a very dark room all make the signal weaker.
  • When it cannot tell — nobody facing the camera, or two people talking over each other — it holds one framing rather than guessing. A frame that jumps to the wrong person is worse than a frame that stays put, so it abstains.
  • It is free. It costs no extra credits either way, so on a two-person recording there is no reason not to tick it.

What it cannot do is invent picture. If both speakers matter at once — an argument, a reaction shot, a demonstration one of them is doing with their hands — then no crop keeps both, and the question stops being about tracking.

The square export exists for exactly that case

1:1 is the compromise nobody asks for and quite often wants. A square crop of a 16:9 frame keeps about 56% of the width, against 32% for vertical. That extra quarter of the frame is frequently the difference between two people in shot and one.

It is also less native everywhere. A square video in a vertical feed leaves bars above and below, which costs you screen. That trade — less height on the phone, in exchange for the second person being visible — is the actual decision, and it depends on what is in the frame rather than on which platform you are posting to.

When to keep the frame you recorded

Sometimes the answer is that this clip is not a vertical clip. 16:9 is the right export when the wide part of the frame is the content:

  • A screen share, a slide, or anything with text on it. Cropped to vertical, the text is off the sides of the frame or too small to read.
  • Gameplay, where the interesting thing is the whole playfield and the minimap is in a corner.
  • A demonstration where somebody's hands and the thing they are holding are at opposite ends of the frame.
  • Anything destined for a landscape player, which includes most embeds and every desktop viewer.

A wide clip with burned-in captions in a vertical feed will look less native than a vertical one. It will also still be legible, which a vertical crop of a slide will not be.

Why there is no page here for each platform

You will find sites with a separate page for TikTok, one for Reels, and one for Shorts. From the export side those three pages describe the same file: a 9:16 video, under a minute, with captions in it. The differences that matter are in what you write in the caption box and when you post, and neither of those is a render setting.

So the choice is not "which platform" but "what is in the frame, and how much of it can I lose". Vertical if the subject is one person or one thing near the centre. Square if two people or a wide subject have to stay in shot. Landscape if the width is the content. If you are unsure, generate the same moment twice — extra candidates from one upload cost a few credits each, not a second full charge for the file.