From a brief to a finished film: how Codex made our predictions video
I gave OpenAI Codex (in the ChatGPT desktop app) the instructions for the video you can watch alongside this post: use our Time Under Tension design system, make it futuristic, exciting and fun, and include an Australian female voiceover with backing music. Codex handled the production from there, developing the story and images, generating the footage and audio, and editing everything together.
My contribution was the brief and creative feedback as the work progressed. I didn’t write the code, operate the video services or assemble the clips, which made this an interesting opportunity to see how much of a connected creative workflow an AI agent could carry out.
The subject was our Collective Predictions, where the Time Under Tension Collective makes its annual predictions about generative AI. We also return to those predictions at year’s end to discuss what we got right and what we didn’t, giving the film a story beyond a collection of impressive futuristic scenes.
Codex began with an outline and moodboard, using our existing brand assets to establish the visual direction. It then generated a starting image for each scene, setting the composition, lighting, materials and subject before adding movement. This gave us something concrete to review early, when changing the look was still relatively straightforward.
Those images went to MiniMax H3 Max through fal.ai, a service that provides access to generative AI models. Codex wrote the code to send each image with instructions describing the action, wait for the generation to finish, and retrieve the resulting video. A still might establish a stylus above a glass interface; the accompanying prompt described the stroke, the interface responding and the camera moving through the shot.
The narration and music came from ElevenLabs, with an Australian female voice generated using its Eleven v4 model and a separate instrumental track prompted to match the film’s energy. The narration also returned timing information for the spoken words, giving Codex a basis for placing captions and cutting between scenes when the story moved to its next idea.
Codex then used FFmpeg, a tool for processing video and audio, to carry out the edit. It trimmed and joined the clips, added our logo and typography, aligned the captions, and lowered the music beneath the narrator. It also added musical accents at scene changes, helping the sound and picture feel connected.
My feedback stayed at the level of the film: faster pacing, more close-up action, futuristic settings throughout and no slow motion. Codex translated those requests into revised prompts and edits, keeping the source assets and production instructions so individual elements could be changed.
What stood out was the amount of coordination sitting behind a simple creative conversation. The workflow required several different services, files and editing steps, yet I could concentrate on whether the film communicated the idea. Codex carried those decisions through production, making the brief and subsequent feedback the main controls for a process that would otherwise require operating each tool separately.