Hey everyone! Lately, I’ve been asked a lot: “Michael, is AI actually ready for commercial cinema and full video production right now?”

Today, I want to share a case study that shatters any skepticism. Director Malik Zenger, who holds Cannes Lions for his music videos but had zero prior AI experience, created a full 22-minute AI film in just 7 days.

Director PJ Ace (@PJaccetturo) broke down this process in detail, conducting an interview with Malik. I have translated and adapted this breakdown for you. This isn’t just dry theory — it’s a step-by-step guide on how a director’s traditional vision, combined with modern tools (like Higgsfield Seedance 2.0 and Soul Cinema), is creating a new cinematic reality. Let’s dive in!)

Here is what the result of this 7-day sprint looks like — the opening scene and teaser for the AI film “Mort”:

Step 1: Narrative & Strategy

Malik Zenger brought a traditional, “old school” filmmaking lens to the project. Instead of chasing “cool AI effects” that flood typical generative videos, he focused on narrative depth and realistic cinematography.

His goal was to prove that modern AI tools can deliver the same emotional weight and drama as a high-budget live-action series like Game of Thrones.

Phase 1: Narrative & Strategy

Step 2: The AI Writer’s Room

Despite having existing visual assets, Malik took the team back to a screenplay development phase. Using Claude as a creative co-pilot, they workshopped a revenge story centering on a complex relationship between two brothers.

This approach turned a one-dimensional, single-use action piece into a full narrative season arc, complete with deep backstory, emotional shifts, and character secrets.

The AI Writer’s Room

Step 3: Organizing the Universe in Figma

To manage the complex Viking world, Malik built a massive interactive board in Figma, organized into “Scene Cards”.

Each card detailed the scene name, the characters involved, and specific location requirements. This visual hierarchy allowed the team to track story beats (such as a child grabbing a wooden sword) to ensure the story’s themes were visually established early on.

Organizing the Universe in Figma

Step 4: Character Architecture & The Face-Blur Hack

The entire character pipeline was anchored in Higgsfield Soul Cinema. This was the key weapon against the cheap, generative “AI look.”

Malik notes that it’s the only model capable of rendering natural human faces with micro-imperfections (pores, wrinkles, a living gaze), completely avoiding the plastic mask look (uncanny valley).

But the biggest production secret was the Face-Blur Hack. Malik discovered that if you leave the faces sharp on wide-shot character reference sheets, the model tries to read microscopic details from a small image and exaggerates them, which results in the final video character having distorted, cartoonish facial features.

The solution: slightly blur the faces on the reference sheet. Then, Soul Cinema focuses perfectly on overall photorealism, maintaining flawless anatomy from any angle, camera setup, or lighting.

Character Architecture & The Face-Blur Hack

Step 5: Extras and The Prop Library

The team detailed reference sheets for secondary characters (e.g., massive Vikings over 2 meters tall) and key props (e.g., the Jarl’s personal emblem).

By defining these assets as character sheets in Soul Cinema, the Seedance 2.0 model recognized them as permanent elements of the universe, keeping the Viking village feeling populated and logically consistent in every scene.

Extras and The Prop Library

Step 6: Location Stress-Testing

Malik noticed that some locations, despite being incredibly beautiful, caused critical bugs in character animation (movements became sluggish or broken).

He generated about six versions of the Viking village and stress-tested them until the character movements in the scene looked as natural as possible. In AI directing, a location is a technical anchor that directly dictates the physics of the actor’s movements.

Location Stress-Testing

Step 7: The Claude Prompt Architect

To translate his directorial vision into a language AI models could understand, Malik built a specific pipeline in Claude. He described the shot artistically and directively (e.g., “handheld, narrow grain, Mel Gibson’s Apocalypto style”), and Claude translated it into a structured technical block for Seedance 2.0.

This guaranteed that every prompt contained clear cinematography terminology (lenses, angles, lighting schemes) before the description of the action. Below is a real example of such a technical prompt for a feast scene in a Viking throne hall, which you can use as a template:

SCENE 93
FRAME 7A
KING’S HALL

— REFERENCE DEFINITIONS —

<<@ JARL>>>: Viking warlord — older now, greying beard longer, deep scar from forehead across left eye to jawline. Heavy fur cloak, chainmail beneath. Weathered, tired. Seated center of the long side of the table facing camera, in his throne — appearance only. Reference.

<<<@ BJORN>>>: Björn the Bull — massive, heavyset Viking warrior, thick neck, broad belly, long dark hair, grey in beard. Ornate golden-brown velvet tunic, fur mantle on shoulders, gold chains on belt. Seated to the right of the Jarl — appearance only. Reference.

<<<@ VIKINGS>>>: Five Viking noblemen — red-faced, drunk. Chainmail, fur, gold rings. Eating, drinking, flirting — appearance only. Reference.

@ women: Three women with full figures — fitted Viking dresses, laced bodices, braided hair, gold chains. Seated among the noblemen, pressed close — appearance only. Reference.

<<<@ KINGSHALL>>>: Viking great hall — vast dark stone chamber. Massive antler chandelier, candles. Large crimson red banner with emblem behind the throne. Cold grey-blue backlight from narrow windows. No fill light. Near-darkness. Long heavy feast table on raised platform with stone steps — the long side of the table faces camera, as seen in reference image. 10 people seated along one side — location and mood reference. Reference.

— TECHNICAL BLOCK —

Flat frontal fill lighting directly from camera position — even, soft, shadowless. Minimal contrast across face and body. No side light, no directional key light, no rim light, no dramatic shadows. The light should feel like a large softbox mounted directly above and around the camera lens, wrapping evenly around all surfaces. Neutral white light temperature.

Shot on ARRI Alexa Mini, Kodak Vision3 250D film emulation, subtle organic film grain. Natural skin with subsurface scattering, visible imperfections — pores, fine lines, uneven skin tone. Fabric texture captured naturally — rough weave, pilling, loose threads. Shallow depth of field with natural optical falloff on close-up view, full sharpness on full-body views. Clean silhouette edges, arms clearly separated from torso in full-body views. Consistent face, body proportions, clothing, and hair across all three angles. No text, no labels, no watermark, no other people, no props in hands, no colored lighting, no ornate decoration. No CGI look, no plastic skin, no airbrushing, no digital smoothing.

The image should feel like a behind-the-scenes wardrobe test photo from a Robert Eggers or Andrei Tarkovsky film — grounded, tactile, lived-in.

— PROMPT —

Wide establishing shot of @ KINGSHALL — vast dark chamber, columns in blackness, antler chandelier above, crimson banner behind the throne. The long feast table faces camera — its long side toward us, all 10 figures seated along it as in the reference. @ JARL center in his throne. @  BJORN to his right. @ VIKINGS and @ women spread along both sides — five noblemen and three women, 10 total. All in the middle of the feast. @ JARL chews slowly, silent. @ BJORN tears meat, mead in his beard. The rest — a tangle of eating and flirting. One nobleman whispers in a woman's ear, she smiles. Another laughs with his head back, goblet raised. A third argues across the table pointing with a drumstick. Two noblemen squeeze a woman between them, arms around her. The third woman is fed a piece of meat, she laughs. Everyone chewing, drinking, touching. The table a mess — roasted meat, bread, gold plates, goblets, spilled mead. In the foreground — a servant crosses frame with a tray, dark and out of focus. Camera holds then slow push-in toward the feast.

SFX only: the feast alive — overlapping laughter, goblets slamming, mead sloshing, meat tearing, chewing, men shouting, women laughing, whispering, plates clinking, drinking horn tilting, servant crossing — soft steps, tray clinking, fading. Candle flames guttering. The hall full of noise.

Step 8: Long Continuous Shots (“Oners” vs. The Cut)

While most AI creators rely on short 2-second clips with fast editing (to hide artifacts), Malik went much further and made things harder: he created 15-second continuous scenes (Oners).

Using precise prompts for “camera shake” and “whip pans,” he achieved a visceral, handheld camera effect. The viewer feels like a direct participant in the action, which tricks the brain into forgetting the entire scene is computer-generated.

Step 9: The Spatial Continuity Hack

How do you keep props (like flipped tables or scattered food) in the same place when changing camera angles? Malik used a “screenshot-as-reference” workflow.

He would generate a wide establishing shot of the village or room, take a screenshot of the resulting “mess,” and upload this frame as a background reference for medium and close-up shots. This allowed him to maintain perfect spatial logic and room geometry, letting the AI freely animate characters within a pre-built space.

The Spatial Continuity Hack

Step 10: Dynamic Stage Progression

Locations weren’t static — they evolved with the storyline. Taking a screenshot of a clean, cozy room, Malik generated a “post-battle version” based on it, with fire, destruction, and chaos.

The resulting frame became the new anchor reference for the next scene. This step-by-step evolution creates a sense of real-time progression and deep narrative.

Dynamic Stage Progression

Step 11: AI Character Voices

Stability and consistency in voiceover were achieved by baking detailed voice descriptions directly into the character cards.

Instead of dry commands, Malik described the voices artistically: “British accent,” or “speaks like a raspy Irish Viking.” The Seedance 2.0 model automatically locked in unique, stable vocal profiles for the characters, preserving the audio track’s integrity throughout the film.

AI Character Voices

Step 12: Real-time Generating & Editing

Malik completely broke away from the traditional “shoot everything first, then edit” mentality.

His work happened in a continuous interactive loop: generate a shot — immediately drop it onto the timeline in the editing program — evaluate lighting and visual cuts with the previous take — generate the next shot accordingly. This approach allowed him to catch any technical discrepancies instantly.

Step 13: The Emotional Secret Sauce — AI-Generated BTS

The most viral and emotional part of the project was the fake “Behind the Scenes” (BTS) featuring animated bloopers.

Malik purposefully generated funny moments: AI actors holding paper coffee cups, someone forgetting lines, or a boom mic accidentally dropping into the frame. Creating this fake reality, where characters behave like “real actors” on set, forged a colossal emotional bond with the audience.


My Conclusions & Takeaways:

This case study once again proves a fundamental truth I constantly repeat to my students: AI tools are just brushes in the hands of an artist.

What makes the film “Mort” feel alive and cinematic is not the algorithm. It’s the director’s traditional vision, deep understanding of shot composition, editing rhythm, drama, and meticulous attention to detail.

When you combine classic laws of filmmaking and drama with a modern AI pipeline, you get a true superpower that lets you shoot full blockbuster-level scenes right from your laptop) Try applying these steps to your own projects!

Thank you for this breakdown and insights! Follow the author on X: @PJaccetturo.