23 short AI films in one afternoon. What 100+ builders and filmmakers taught us about running video models under real deadlines, and what we changed because of it.
September 17, 2026

On September 13 we co-sponsored the AI Filmmaking Masterclass + Hackathon at Yes SF in San Francisco, hosted with The Multimodal Society. 195 people applied. The ones who got in spent the morning learning how to direct a video model like a director, then spent four hours in teams making a short film from scratch. At 5:30 PM we screened all 23 of them. The 3 hours masterclass was lead by Roan Weigert.

Every team had credits from event partners: Magnific, Tripo AI for 3D generation, ElevenLabs for voice and sound, and GMI Cloud for image and video inference. Teams could call models through the GMI Cloud inference engine directly or through GMI MCP from inside the agent tools they already used.
Two tracks. Creative, for films. Developer, for tools and interactive experiences built on top of the models. Judges from OpenAI, ByteDance, and Alumni Ventures scored story, execution, originality, and impact, with a bonus for how well teams used the partner stack. Judges:
Long Do, Member of Technical Staff, OpenAI
Calissa Man, Multimodal Post-training, ByteDance
Bryan Chellen, Principal, Alumni Ventures

A hackathon compresses months of typical customer behavior into an afternoon. A few patterns showed up fast.

Iteration count beats prompt length. Teams that wrote long, careful prompts and waited produced fewer usable shots than teams that fired short prompts, looked, adjusted and fired again. Latency per generation set the ceiling on how many tries a team got. That is the metric that decided the afternoon, more than any quality benchmark.
Consistency is the hard problem. Almost every team hit it: the same character, the same product, the same look across dozens of shots. The teams that finished treated consistency as a workflow, with reference images, a locked style, and a rule for when to give up on generation and cut to a real photo. The ones that tried to fix it prompt by prompt ran out of time.
Mixed media wins. The films that landed with the audience blended generated scenes with recorded footage, motion graphics and real voice. Pure generation looked like pure generation. The models are a department in the production, and a good producer knows which department to send each shot to.
Tool selection is a skill. The most requested topic in the applications was which tool to use for which job. We spent forty minutes of the masterclass on it and could have spent two hours. Image model for the key frame, video model for motion, 3D for objects that need to hold shape across angles, upscaler last. Developers who internalized that order shipped.

Developer track: Prison Escape! An Interactive Video Experience, by Peter Martin. A branching film where the audience picks the escape route. It ranked first overall across all judges and took the top impact score of the day. It is also a preview of a format that becomes possible when generation is fast and cheap enough to render every branch.

Creative track: In Her Language, by James Yin and team, with Xiaoqing Lin on story. Highest story and originality scores. Proof that the emotional ceiling on AI film is set by the writer, and by nothing else.

A nine-year-old also presented his film, Gut Run, to the full room and won a nominated award. He handled questions better than most founders I have filmed.

Best use of TripoAI, ARf-ARf: Video link

Ranked by the judges’ average where scores were complete.
Developer track
Creative track
In Her Language, winner
Gut Run, nominated award
Watching people work under a deadline is the fastest product research there is. Three things went straight into our backlog.
First, per-generation latency is a creative constraint, and we should surface it as one. A filmmaker deciding between two models cares about how many tries they get before lunch ends.
Second, consistency workflows deserve first-class support in the tooling. Reference handling and style locking came up in nearly every team conversation, and people were solving it with folders and sticky notes.
Third, agent-native access matters. Teams that called models through GMI MCP from inside their coding agents stayed in one window and iterated faster than teams switching between browser tabs. Video generation is becoming a tool call inside a larger workflow, and the developer experience needs to reflect that.
To Liz Zhang and The Multimodal Society for co-hosting, Loc Nguyen and Bayhaus Creative for production, Yes SF for the space, and Yuqi Hou from our community team for running the GMI Cloud table all day. And to the 23 teams. You made more finished work in four hours than most production schedules allow in a month.
More dates are coming, San Francisco first. If you build on video models and want to see how filmmakers use your work under pressure, this room is the place to be. Follow GMI Cloud on LinkedIn or reach me directly for the next date.
If you want to try the same stack the teams used, the GMI Cloud inference engine and GMI MCP are available today at gmicloud.ai.
Roan Weigert
DevRel Lead @ GMI Cloud
GMI Cloud helps you architect, deploy, optimize, and scale your AI strategies
