Field Report · August 2026

Do Filmmakers Need an Agentic AI, or Just a Faster Assistant?

Do Filmmakers Need an Agentic AI, or Just a Faster Assistant?

A field report on invideo Agent Two — and the question it raises for every creative director watching AI video mature in real time.

The Brief

Few day ago I ran a real experiment, not a demo. One STORY. No reference images, no style frames, no shot list:

Create a 120-second sci-fi mockumentary titled "Welcome to Earth 3026," shot in the style of a prestige streaming documentary (think BBC/Netflix future-history specials) — sober narrator voice-over, drone establishing shots, archival style intercuts, infographic overlays, and reflective talking-head interview cutaways. No action, no violence, no dialogue scenes — this is journalistic, calm, and analytical in tone.

Open in the present tense of the year 3026: Earth is a clean, green, thriving civilization. The United Nations has become the single most powerful governing body on the planet, with representatives from every nation's government sitting as one council. There has been no war, no extreme disease or pandemic, no poverty, and no money for generations. Crime is virtually nonexistent because repeat bad actors are exiled to outer-planet colonies rather than imprisoned.

Society runs on a Contribution Credit system: every person is rewarded with credits based on how much they contribute to society, and those credits are redeemed for better housing, vehicles, clothing, and personal AI assistants. Show this visually through elegant infographics and everyday citizens receiving upgraded living support.

Midway through, flash back in archival-newsreel style (grainier, desaturated, multi-language chaotic news broadcasts) to the year 2222: a small meteor struck Earth and spread a deadly disease, killing one billion people within a single year before scientists found a cure. Show the global panic, the scientific breakthrough, and the moment nations set aside their conflicts and united — this catastrophe is the origin of the current era of peace.

Close with reflective interview-style commentary on how that shared tragedy forged permanent unity, then end on humanity's common goals for the future: protecting life on Earth from any threat, internal or external, and exploring other planets and galaxies. Final shot: a deep-space vessel launching outward as the UN insignia dissolves into stars, with the title card "WELCOME TO EARTH 3026.

Documentary color grading throughout — natural but slightly idealized lighting for the 3026 present, desaturated archival grain for the 2222 flashback. Steady, contemplative pacing. Generate at 720p.

That's it. That's all Agent Two got to start.

What came back, roughly 90 minutes and 187.5 credits later, was a fully scripted, storyboarded, narrated, shot, and assembled 120-second film — built by an AI agent that didn't just generate clips, but directed them.

What "Directing an AI Crew" Actually Looked Like

Agent Two didn't ask me for a shot list. It asked me three creative-direction questions — aspect ratio, narrator voice, and a verbatim closing line — then went to work:

  • Wrote its own 13-beat AV script, down to the second, with narration and shot direction for every beat
  • Storyboarded three acts as still frames before spending a single video-generation credit — so I could catch problems early
  • Cast its own narrator, auditioning four AI voice candidates and recommending one
  • Generated 24 video shots, kept 17, and self-corrected along the way
  • Caught its own mistakes — a background shot that accidentally showed recognizable real-world leaders (which I asked it to fix), a shot that stubbornly kept rendering readable text in a supposedly moneyless society (which it eventually cut rather than keep spending credits on), a dropped final act in the first render, and a resolution mismatch — all diagnosed and fixed without me digging through timelines myself

I answered questions. I did not touch a script editor, a shot list, or a video timeline until the very end.

The Honest Numbers

  • Total cost: 187.5 credits
  • Total time: ~91 minutes, script to final master
  • Image generation hit rate: 4 used / 5 generated (80%)
  • Video generation hit rate: 17 used / 24 generated (71%)
  • Single biggest cost center: one stubborn shot that ate roughly 12% of the entire budget before being abandoned

That last number matters. Agent Two's instinct is to keep retrying a problem shot rather than cut it early — a real workflow lesson if you're managing a credit budget. It got there eventually, but not before spending real money finding out.

Where It Genuinely Impressed Me

The self-correction was the standout. Not one issue in this project was something I had to catch by scrubbing through footage. Agent Two flagged its own render failures, explained what broke and why, and fixed it — narrating its own QA process in plain language. That's a meaningfully different experience from directing a single-shot generator one prompt at a time.

Where It Still Needs a Human

It still didn't hold my 720p instruction cleanly on the first export pass — I had to catch that at the finish line. And even with a beautiful assembly, I still opened video editing software afterward for more complex editing that any AI Agent could not handle. Creating opening title animation, adding subtitle, adjusting speed for certain shot, adjusting color, removing glitch inside AI video generation and other final polish — that's still a human job. Agent Two builds an excellent base. It doesn't yet replace a final creative pass.

So, Agent or Assistant?

Here's the distinction I keep coming back to. A faster assistant executes what you tell it, one instruction at a time — you're still the one catching every mistake, re-rendering every bad shot, holding the whole timeline in your head. An agent manages the process — it catches its own dropped act, argues with its own stubborn overlay, and hands you a finished cut to react to instead of a pile of raw clips to assemble.

Agent Two behaved like the second one, more often than not. Not perfectly — I still had to step in on real-world lookalikes and a resolution slip — but the direction of travel is clear: less "faster tool," more "junior director you're mentoring."

Whether that's what your workflow actually needs is a different question entirely.

The Question I Actually Want to Ask You

This is the part I'm most curious about, because I don't think there's one right answer yet.

Given where AI video tools are today — would you rather generate video through a manual, shot-by-shot workflow where you control every prompt and every model choice, or an agentic workflow where you direct a crew of AI agents and review their output at checkpoints?

Manual gives you precision and full creative control, at the cost of your time. Agentic gives you speed and self-correction, at the cost of some control you have to claw back at review points.

I've now worked both ways extensively. I still don't think either one is simply "better" — they solve different problems. But I want to hear where the room stands.

Drop a comment: manual, or agentic — and why?


Back to Blog