
Runway is an AI video platform built around generative world models β Gen-4.5, Aleph 2.0, and Act-Two for performance capture β split across Runway Creative for filmmakers, Runway Dev for builders, and Runway Robotics.

Loadingβ¦

ShortMatic AI turns long videos into vertical short-form clips in the browser. It finds high-engagement moments rather than cutting on a timer, adds word-timed captions, and reframes around faces for 9:16, 4:5 and 1:1.
Turning a podcast into short clips is three separate jobs, and only one of them is hard. Cutting is trivial. Captioning is solved. Deciding which forty seconds of a ninety-minute conversation is worth posting is the job that normally eats an editor's afternoon.
ShortMatic AI is built around that third job. Its Moment Finder reads the whole video and picks segments it judges high-engagement, then captions and reframes them for vertical platforms β all in a browser, from a YouTube link or an upload.
This review covers what moment detection can genuinely infer, why word-timed captions are not a cosmetic feature, where auto-reframe breaks, and who is better served doing this by hand.
This is the question the whole category turns on, and the honest answer is: some of it, reliably, and the rest not at all.
What a model can detect well is structural. A question followed by a pause and a long answer. A change in vocal energy. A laugh. A phrase repeated for emphasis. A speaker interrupting themselves. These are real signals of something happening, they are visible in audio and transcript, and picking them beats cutting on a timer by a wide margin.
What it cannot judge is whether the moment means anything to your audience. A clip's performance depends on context the video does not contain β what your viewers already believe, what is being argued about this week, which of your guest's opinions is the controversial one. A model reading the file has none of that.
The right expectation, then, is triage rather than selection. It hands you twelve candidates from ninety minutes, which is a genuinely useful reduction. Choosing the three to post is still yours, and that is where the judgement lives.
It can find where the energy changed. It cannot know which sentence your audience will argue about in the comments.
Word-timed captions, offered as Bold Reels, One Word, or none. Worth understanding why this matters more than the styling options suggest.
Short-form video is overwhelmingly watched without sound. A clip with no captions is not a quieter clip β for most of its audience it is a silent one, and it will be scrolled past regardless of how good the moment was.
Word-timed matters specifically. Captions that appear a phrase at a time force the viewer to wait; captions that highlight each word as it is spoken keep the eye moving with the audio, which is the mechanic behind the style that dominates the format.
The One Word style β a single word filling the frame β is the aggressive end of that. It commands attention and it suits a punchy quote far better than a nuanced explanation. Bold Reels is the safer default for anything where the sentence matters more than the impact.
Auto-reframe crops subject-aware into 9:16, 4:5 and 1:1, tracking faces and action rather than trusting the centre of the frame.
That distinction is not academic. A two-person interview shot wide has both speakers well off-centre, and a naive vertical crop produces a clip of the space between them. Face tracking is what makes the format conversion survivable at all.
It breaks in predictable places, and knowing them saves a wasted upload. Fast cuts between speakers can make the crop chase the frame. Anything where the meaning is in the wide shot β a whiteboard, a screen share, a demonstration using both hands β loses the point when cropped to a face. Three or more people on screen force the crop to choose, and it will sometimes choose wrong.
The practical rule: talking heads reframe well, and anything visual does not. If your long video is a screen-recorded tutorial, vertical conversion is the wrong idea regardless of which tool does it.
The free plan needs no card, so the only sensible evaluation is on your own footage rather than a demo reel.
If your long-form content is visual rather than spoken, this category does not fit you. A tutorial where the value is on the screen cannot be cropped to a face without discarding the point.
If you post rarely, the setup and review time is unlikely to beat clipping by hand once a month.
And if clip selection is your craft β if the reason your shorts work is that you know exactly which line lands β then an automated shortlist is at best a second opinion. Some creators genuinely are the algorithm, and for them this saves the easy part while touching none of the hard part.
It is also worth noting this is a small, young product. The site is browser-based with a free tier, a blog and a FAQ, and it publishes terms, privacy and refund pages. Keep your source files; that is good practice with any hosted editor, not a comment on this one.
These all cut long video into short, and differ on what else they try to own.
| Tool | Core job | Moment selection | Where it runs |
|---|---|---|---|
| ShortMatic AI | Long video to captioned vertical clips | Moment Finder over the whole video | Browser, YouTube link or upload |
| Munch | Clipping with marketing analytics attached | Yes, with trend signals | Browser |
| Quso.ai | Clipping plus wider social workflow | Yes | Browser |
| SendShort | Short-form generation and scheduling | Yes | Browser |
| Zeemo | Captions and subtitles specifically | No β you choose the clip | Browser and app |
| Category | AI short-form video clipping |
|---|---|
| Named features | Moment Finder, Captions, Auto-reframe |
| Inputs | YouTube links, or uploaded long-form files |
| Typical source material | Podcasts, tutorials, interviews, webinars |
| Aspect ratios | 9:16, 4:5, 1:1 |
| Caption styles | Bold Reels, One Word, or none |
| Caption timing | Word-timed |
| Reframing | Subject-aware, tracking faces and action |
| Export targets | YouTube Shorts, TikTok, Instagram Reels, Facebook Reels |
| Runs in | The browser, no installation |
| Free plan | Yes, no card required |
| Stated audiences | YouTubers, podcasters, educators, agencies |
| Site sections | Features, how it works, blog, FAQ, terms, privacy, refunds |
Clipping, captioning and assembly are separable β most creators end up with two of the three.
Last checked . Taken from shortmatic.app during this check, with the screenshot above captured at the same time.
Observations about what moment detection can and cannot infer, and where subject-aware reframing breaks, describe the technique in general rather than measured behaviour of this product. Clip performance was not tested.
Moment Finder, analysing a whole video to pick high-engagement segments
YouTube link input as well as direct file upload
Word-timed captions in Bold Reels or One Word styles
Subject-aware auto-reframe tracking faces and action
Output at 9:16, 4:5 and 1:1
Export ready for YouTube Shorts, TikTok, Instagram Reels and Facebook Reels
Runs entirely in the browser with no installation
Free plan with no credit card required
AInexfinder does not list plan prices or billing details. For current pricing, plans, and trials, visit the official ShortMatic Ai website.
Visit official website for pricingVendor pricing, credits, and billing policies change over time. Always confirm on the official site before you buy.
Log in to write a review.
No reviews yet. Be the first to share your experience with ShortMatic Ai.
Assigned reviewer
Daniel ReedSenior AI Tools Reviewer
Daniel reviews AI tools the slow way β by actually using them on real projects. His reviews cover what works, what breaks, and who each tool is genuinely a good fit for.
Daniel and the AInexfinder editorial team research ShortMatic Ai using public product information, listing evidence, and (when available) hands-on checks. Scores reflect listing completeness and transparency β not paid placement.