8 AI video editing prompts for a better first cut
Use AI as an edit-planning assistant, not as proof that the edit is finished. Paste a timestamped transcript or shot log, name one job, set the output format, and tell it what it must not invent. These eight prompts cover a rough cut, opening, captions, B-roll, pacing, transitions, color notes, and repurposed clips. Check every suggestion against the footage before changing the timeline.
01 · SOURCE BEFORE PROMPT
What should you give an AI before asking it to edit?
A useful editing prompt starts with evidence the model can point back to. OpenAI recommends clear, specific instructions with enough context, while Anthropic recommends stating the desired format and constraints. For an edit plan, that context is usually a timestamped transcript, a shot inventory, or both.
OpenAI: prompt engineering best practices for ChatGPTAnthropic: prompting best practices
| Input | What to include | What it prevents |
|---|---|---|
| Timestamped transcript | Exact words, speaker changes, and uncertain text. | Invented quotes and vague cut points. |
| Shot log | File names, time ranges, framing, and usable moments. | Suggestions for footage that was never recorded. |
| Editing job | One goal such as rough cut, captions, or B-roll. | A generic answer that mixes unrelated decisions. |
| Output schema | Named columns, labels, or a file-shaped draft. | Advice that cannot be transferred to the timeline. |
| Guardrails | Facts, claims, footage, rights, and context to preserve. | A cleaner edit that changes what actually happened. |
02 · COPY THE JOB, NOT THE HYPE
8 AI video editing prompts to copy
Replace every bracketed field. Keep file names and timestamps intact. Each prompt asks for a plan or draft, never a claim that the AI touched your media. If your assistant can inspect images or video, give it only the assets you have permission to share.
You are planning a rough cut, not editing the media files. INPUTS - Target viewer: [who] - Target runtime: [length] - Required facts or moments: [list] - Timestamped transcript: [paste] - Shot log with file names and time ranges: [paste] TASK Build a rough-cut decision table with these columns: 1. source time range 2. KEEP, CUT, MOVE, or VERIFY 3. exact spoken line or shot description 4. proposed timeline position 5. reason tied to clarity, continuity, or the stated goal 6. editor action RULES - Use only words and shots present in the inputs. - Never invent a quote, reaction, camera angle, result, or piece of B-roll. - Mark uncertain transcript text or missing visual evidence as VERIFY. - Preserve required facts even when they slow the pace. - Do not claim that any cut has been applied or exported. Finish with the proposed runtime and a short list of unresolved checks.
Give it: A timestamped transcript, shot log, target length, and must-keep facts. Get back: A source-to-timeline table with keep, cut, move, and verify decisions. Why this works: This makes the model show where every recommendation came from. You can reject one row without rebuilding the whole plan.
Act as an opening editor. I need three truthful options for the first 8 seconds. INPUTS - Intended viewer: [who] - Verified payoff in this video: [what the footage really shows] - Timestamped transcript: [paste] - Shot log: [paste] FOR EACH OPTION, RETURN - exact in and out source timestamps - verbatim spoken words, in order - the first visible shot - the question or promise created - what later moment pays it off - any continuity risk RULES - Do not rewrite spoken words and present them as recorded audio. - If a new line would help, label it PICKUP TO RECORD and keep it separate. - Do not imply a result, test, or personal experience missing from the footage. - Do not choose a visual that is absent from the shot log. - Do not claim one opening will increase retention. Rank the options by clarity, not predicted virality.
Give it: The transcript, shot log, verified payoff, and intended viewer. Get back: Three openings with exact source references and any pickup line labeled. Why this works: A stronger opening is useful only if the footage can deliver it. Source references stop a catchy rewrite from becoming a false promise.
Turn the checked transcript below into an SRT-shaped caption draft. INPUTS - Language: [language] - Speaker names and spelling glossary: [paste] - Reading constraints: [maximum lines and characters, if known] - Checked transcript with timestamps: [paste] TASK - Split captions at natural phrase boundaries. - Keep the original meaning and wording. - Use sequential SRT numbers and HH:MM:SS,mmm timestamps. - Put uncertain words in [[double brackets]]. - After the draft, list every name, number, technical term, overlap, and noisy section that needs review. RULES - Do not invent missing speech or translate unless I explicitly ask. - Do not paraphrase a claim to make it shorter without marking the change. - Do not guess speaker identity. - Do not claim the draft is ready to import. Return only the SRT-shaped draft and the verification list.
Give it: A checked transcript, names glossary, language, and caption rules. Get back: An SRT-shaped draft plus a list of words and timestamps to verify. Why this works: Caption formatting is mechanical, but names, accents, and timing still need a human pass before import or publication.
Create a B-roll plan for this edit using the footage inventory as the source of truth. INPUTS - Timestamped transcript: [paste] - Available footage and stills, with file names and time ranges: [paste] - Rights or disclosure notes: [paste] - Visual style constraints: [paste] RETURN A TABLE WITH 1. timeline moment 2. spoken idea the visual must support 3. available asset and exact source range 4. framing or crop note 5. duration estimate 6. status: AVAILABLE, VERIFY RIGHTS, or NEW SHOT NEEDED RULES - Never invent an available asset. - Use an asset only when its description supports the spoken idea. - Label every missing visual as NEW SHOT NEEDED; describe the shot, but do not claim it exists. - Keep licensed, sponsored, or AI-generated material visibly flagged for review. - Do not claim that any media was placed on the timeline. End with the three highest-priority coverage gaps.
Give it: The transcript, available asset inventory, and rights notes. Get back: A B-roll plan that separates usable footage from new shots to capture. Why this works: The prompt treats missing coverage as a production decision instead of silently inventing an asset that the editor cannot find.
Review this cut plan for pacing risks. Treat every conclusion as an editing hypothesis. INPUTS - Video goal and intended viewer: [paste] - Timestamped transcript: [paste] - Current cut or edit decision list: [paste] - Known shot, text, or topic changes: [paste] RETURN A TABLE WITH 1. timeline interval 2. observable issue in the supplied material 3. why it may slow comprehension or create overload 4. one edit to test 5. information or continuity that edit could damage 6. verification needed RULES - Do not infer audience retention, views, or algorithm response from the transcript. - Do not mark a pause as dead air when it serves a demonstration, emotion, or comprehension. - Do not remove required context to make the cut faster. - Use only the supplied timing and shot-change data. - Do not claim the edit has been made. Finish with one conservative pacing pass and one more aggressive pass.
Give it: A timestamped transcript, current cut notes, and known shot changes. Get back: Intervals to inspect, one proposed action each, and a stated tradeoff. Why this works: Pacing is not the same as constant speed. The output protects necessary context and keeps every diagnosis as a hypothesis until data or review supports it.
Create a transition plan for the proposed cut below. INPUTS - Ordered rough-cut table: [paste] - Audio notes, room tone, and music limits: [paste] - Any intentional changes in time, place, speaker, or topic: [paste] FOR EACH CUT, RETURN - outgoing and incoming source ranges - relationship: CONTINUOUS, NEW TIME, NEW PLACE, NEW TOPIC, or CONTRAST - recommended transition: HARD CUT, J-CUT, L-CUT, DISSOLVE, or HOLD - why the transition clarifies that relationship - audio or continuity check RULES - Default to a hard cut when no other transition has a clear job. - Do not add sound, room tone, music, or footage that is not listed. - Do not hide a continuity error with a transition; flag it. - Do not use a dissolve to imply time passing unless the story supports that meaning. - Do not claim the transitions were applied. End with the cuts that require an editor's judgment in playback.
Give it: The rough-cut order, audio continuity, and relationship between adjacent shots. Get back: A transition map with a default hard cut and exceptions justified. Why this works: Transitions should communicate time, place, continuity, or contrast. Starting with a hard cut avoids decorating every edit by default.
Act as a color-review assistant. Write correction notes, not a finished grade. INPUTS - Reference frames with shot IDs: [attach or describe each visible frame] - Camera, profile, and lighting notes if known: [paste] - Intended visual reference: [attach or link] - Delivery platform and color space if known: [paste] RETURN FOR EACH SHOT - visible mismatch relative to the supplied reference - correction goal for exposure, white balance, contrast, and saturation - skin-tone or product-color check when relevant - what must be measured on scopes - status: OBSERVED IN FRAME, PROVIDED METADATA, or ASSUMPTION TO VERIFY RULES - Do not invent camera settings, color temperature, exposure values, LUTs, or scope readings. - If no frame is supplied, say that visual assessment is not possible. - Separate technical normalization from creative look choices. - Preserve deliberate lighting differences unless I identify them as errors. - Do not claim that a grade has been applied or exported. Finish with a shot-matching order that starts from the best reference frame.
Give it: Reference frames, shot IDs, camera or lighting notes, and delivery target. Get back: Correction goals per shot, with measurements or assumptions flagged. Why this works: Text alone cannot reveal image data. Requiring frames and labeling assumptions keeps the plan useful without inventing exposure or color values.
Find up to three self-contained short clips in this longer video. INPUTS - Timestamped transcript: [paste] - Shot log: [paste] - Target platform and maximum duration: [paste] - Facts, disclosures, or context that must stay attached: [paste] FOR EACH CLIP, RETURN - exact source in and out timestamps - verbatim first line and final line - the single question or takeaway it can answer alone - required context inside the clip - matching available visuals - estimated source duration - any PICKUP TO RECORD, clearly separated from existing speech RULES - Keep the speaker's meaning intact. - Do not combine nonadjacent words into a sentence that was never spoken. - Do not remove a qualification, disclosure, or condition that changes the claim. - Do not invent B-roll, performance results, or a title presented as proven. - Do not claim the clips were cut or published. Reject any candidate that cannot stand alone within the stated duration.
Give it: A timestamped transcript, shot log, platform limit, and required context. Get back: Three clip candidates with exact boundaries and any pickup need labeled. Why this works: The model must find complete ideas already recorded, not manufacture a hook that changes the speaker's claim.
03 · MOVE FROM PLAN TO TIMELINE
How do you turn an AI answer into a real edit?
Treat the answer as an edit decision list. It can reduce the blank-page work, but the footage, playback, and final timeline remain the authority. A five-step pass catches most failures.
- 01
Trace every note to a source
Open the named file and source range. Reject any quote, shot, or timing decision that cannot be found.
- 02
Apply a duplicate or a short test range
Keep the original sequence. Try the highest-impact cuts on a duplicate before changing the full edit.
- 03
Watch for continuity and changed meaning
Check the audio, eyeline, action, disclosures, and every qualification that makes a claim accurate.
- 04
Validate captions and generated assets separately
Review spelling and timing before import. Confirm rights and disclosure needs before using generated footage.
- 05
Compare the published result with real evidence
Use playback review, comments, and platform analytics to test the hypothesis. Do not backfill a causal story from one view count.
If you are deciding whether this workflow is worth the time and subscriptions, use Kabo's guide to AI video editing cost. If a platform has already assembled the first cut, the checklist for reviewing Instagram First Draft shows the same source-first discipline in a product workflow.
04 · MATCH THE TOOL TO THE JOB
Which AI tool should you use?
Use a text assistant for a transcript-based edit plan. Use a multimodal assistant only when you can attach frames or footage safely and it can cite the supplied material. Use a video generator when the job is to create a candidate clip, and use an editor when the job is to change the timeline and export. Those are different actions.
Google Flow, for example, can generate or extend clips from text, frames, or other ingredients. That does not mean a clip has been reviewed, licensed for your use, placed into your existing sequence, or published. Keep generation and editing as separate checklist items.
Google Flow Help: create videos in Flow
| Tool type | Useful for | Still needs verification |
|---|---|---|
| Text assistant | Transcript, caption, and decision-list drafts. | Source accuracy, timing, visuals, and final playback. |
| Multimodal assistant | Reviewing supplied frames or clips with text context. | What was actually inspected and what remains unseen. |
| Video generator | Creating a candidate shot when generation is the job. | Rights, disclosure, continuity, quality, and placement. |
| Video editor | Timeline changes, playback, captions, and export. | Meaning, claims, rights, and audience evidence. |
05 · KEEP THE CLAIM SMALL
What can these prompts not verify?
The prompts cannot see footage you did not provide. They cannot prove that an edit caused retention, views, recommendations, or sales. They cannot establish rights to a clip, identify a person, confirm a disclosure, or know whether a line was recorded truthfully. They also cannot promise that an SRT-shaped answer will import without errors.
YouTube's audience-retention report describes moments where viewers start or stop watching, but the report still needs context and is not instant. YouTube says the data typically takes one to two days to process. Use dips and spikes to decide what to inspect, then compare the video, traffic, comments, and your edit notes before forming a hypothesis.
06 · FAQ
Common questions about AI video editing prompts
Can ChatGPT or Claude edit my video from a prompt?
A chat assistant can plan edits from material you provide. Whether it can inspect media or change a timeline depends on the product and tools connected to it. Keep the prompt's output labeled as a plan until you verify a real timeline.
What if my transcript has no timestamps?
Ask for structural notes by paragraph or sentence, not precise cut points. Add timestamps before requesting an edit decision list. Automatic captions can be a starting point, but correct uncertain words first.
Can I ask AI to make an SRT file?
You can request an SRT-shaped draft. YouTube supports basic SRT files as plain UTF-8 text, but you still need to validate the numbering, timestamps, text, and import in your editor or platform.
Do these prompts work for long and short videos?
Yes, if you replace the runtime, platform, and context constraints. Do not force a long explanation into a short clip by removing a qualification that changes the claim.
07 · SOURCE LEDGER
Sources and method
Product and platform facts were checked on September 16, 2026. The official sources support prompting principles, video generation boundaries, caption review, SRT formatting, and retention-report behavior. The eight prompts, input table, and five-step verification pass are Kabo editorial material. They are not product guarantees, tested performance results, or a claim that an AI system completed an edit.