Vertical clips with captions
Self-contained moments are found automatically, cut with a short pre-roll, and captioned word-for-word from the transcript. Cropped to 9:16, 1:1, or 4:5 depending on where the clip is headed.
Most teams record more than they publish. Tocamoufy takes one long recording — a talk, a podcast episode, a webinar — and produces the formats that usually get skipped: vertical clips, a written summary, pull-quote graphics, a thread outline, timestamped show notes, and chapter markers. Everything lands in a queue for review before anything exports.
Clip selection, captions, the summary, and the show notes all read off one aligned transcript and timing pass over the source file, so they stay consistent with each other. Each output arrives as an editable draft, not a finished asset.
Self-contained moments are found automatically, cut with a short pre-roll, and captioned word-for-word from the transcript. Cropped to 9:16, 1:1, or 4:5 depending on where the clip is headed.
Organized around the topics the speaker actually returned to, with headings that follow the source's own section breaks rather than a generic outline.
Sentences with a clear claim or sharp phrasing are flagged as candidates, then laid out on a fixed-size card in one of three included styles.
The summary's structure reformatted into numbered posts sized to a standard character limit.
Every section break and clip becomes an HH:MM:SS line, formatted for a podcast host or a video description field.
The same timestamps written back out as a chapter file for the original long-form upload.
Tocamoufy starts by transcribing the source and aligning that transcript to the audio at the word level — the same alignment step that drives captioning later. Clip selection then runs against the transcript, not the raw waveform: it looks for spans that read as complete on their own, a question and its answer, a claim followed by a specific example, an anecdote with a beginning and an end, and scores each one for length, pacing, and how much it depends on something said earlier in the recording. A span that leans on unresolved context loses points even when the delivery is good.
Where there's more than one speaker, the system tracks speaker turns so a cut point doesn't land mid-interruption. For webinars and screen recordings, it also watches for slide changes and treats them as natural boundaries, which tends to produce cleaner cuts than audio energy alone would.
What comes out of this step is a ranked list of candidates — usually two to three times more than you'd actually use — not a fixed set that gets exported automatically. You keep, reorder, or throw out anything from that list once it reaches the editing queue.
Nothing Tocamoufy generates gets posted anywhere. Each candidate clip, quote card, or note lands in an editing queue as a draft, and exporting it is a separate, manual step you trigger per asset.
In the queue, a clip's in and out points are draggable handles, not fixed values — extend one to keep a laugh in, or trim it if the setup ran long. Captions are editable text, line by line, generated from the same transcript that drove clip selection, so fixing a misheard word doesn't require reprocessing the source. Reframing for a vertical or square crop starts from a suggested focus point, usually the active speaker, that you can override by dragging the crop window yourself, per clip, per aspect ratio.
Pull-quote cards and the written summary open in a plain-text editor before export, and the thread outline is just text you can rewrite outright. Nothing in the process assumes the first draft is final — it's built around the expectation that you'll change several things before anything leaves the queue.
No auto-posting. Tocamoufy doesn't connect to social accounts. Every export is a file — video, image, or text — that you download and post yourself, wherever and whenever you decide to.
What goes in, what comes out, and the limits on both.
| Property | Value |
|---|---|
| Container / file type | MP4, MOV, MKV, WAV, MP3, M4A |
| Max length per upload | 4 hours |
| Max file size | 8 GB |
| Video resolution | up to 3840×2160, downsampled to 1080p for processing |
| Frame rate | 23.976–60 fps, variable frame rate accepted |
| Audio channels | mono or stereo; per-speaker multi-track isolation not yet supported |
| Minimum audio quality | intelligible speech at 96 kbps or better |
| Output | File type | Dimensions | Length / size | Captions |
|---|---|---|---|---|
| Vertical clip | MP4 (H.264) | 1080×1920, 1080×1080, or 1080×1350 | up to 3 min / clip | burned-in or .srt |
| Pull-quote graphic | PNG or SVG | 1600×2000 | — | n/a |
| Written summary | Markdown or .docx | — | roughly 400–900 words | n/a |
| Thread outline | Markdown or .txt | — | sized to a 280-char/post default | n/a |
| Show notes | Markdown or .txt | — | one line per timestamp | n/a |
| Chapter file | .txt (YouTube-style) or .xml | — | — | n/a |
Every plan includes all six output types. The difference is how many hours of source recording you can process each month and how quickly it gets processed.
Single
$24/month
For one recurring show or a monthly recording, run by one person.
Studio
$69/month
For a small team publishing from a few recordings a week.
Team
$179/month
For teams running multiple shows or a steady conference/webinar schedule.
Billed monthly, no annual lock-in. Additional hours beyond a plan's monthly allotment are billed per hour processed.
No. Tocamoufy produces drafts — clips, graphics, and text files — that you download or export manually. It doesn't connect to social accounts and has no path to publish on your behalf.
The uploaded file and its transcript are kept for 30 days after upload so you can return to the editing queue, then deleted automatically. You can delete a project sooner from inside it.
The transcript is speaker-labeled, and clip selection uses those speaker turns to avoid cutting in the middle of an interruption. If a label gets a name wrong, it's editable directly in the transcript view.
Caption presets currently cover a handful of built-in styles — size, position, and one or two accent colors. Custom font uploads aren't supported yet.
Clip selection scores for pacing as well as content, so long pauses tend to score low and rarely surface as candidates. It won't recognize a technical failure as unusable on its own if the audio is otherwise clean — that judgment call stays with you in the queue.