b07832d1-dd67-481d-8899-225d6674cbed MODEL
MAKER
d96d2406-8ab9-49e1-a95e-2f4d0e6a7902
ACCESS
api
TRAINABLE
Yes
GEOMETRY
/ 5
Ask most video models for twenty seconds inside a courtyard and you get a camera move. Ask FLUX 3 and the footsteps arrive with the frames, generated in the same pass, landing roughly where the feet do. Black Forest Labs treats that timing as evidence rather than as an audio feature: a model that knows what an event sounds like has learned something about the event that a model trained only on pictures never had to.
That argument is the reason to pay attention to FLUX 3, more than the clips it produces this month.
FLUX 3 is Black Forest Labs' frontier model and its first video model. The company is German, founded by researchers who worked on Stable Diffusion, and it built its reputation on stills: FLUX.1 in August 2024, then the FLUX.2 family. FLUX 3 went into early access in July 2026, and video generation reached general availability on 4 August 2026.
It is multimodal, meaning it was trained on video, images and audio together inside one architecture rather than as separate systems wired to each other after the fact. You can start from text, from a still, from a start and end frame, or from an existing clip. Clips run up to twenty seconds in a single generation, at 24fps, in HD or Full HD, with optional audio — dialogue, effects and ambience — generated alongside the frames at no extra charge.
The entry point for most practices is not a text prompt. It is the hero render you already have. Feed a still in and the model extends it into motion; give it a start and end frame and it fills the interval. That is the closest thing to control on offer.
On Black Forest Labs' published rates, standard HD runs $0.17 per second, putting a twenty-second clip at $3.40. A draft pass — a fast, rough preview you approve before committing — costs $0.06 per second, about $1.20 for the same clip. Full HD is $0.29. [verify: rates current at publication; partner platforms mark up or convert to credits]
Those numbers change one decision, and it is a small-practice decision: whether moving imagery goes into a competition submission at all. A mood film for a bid, a sequence storyboarded before you commission the real animation — work that previously had to be bought or skipped. None of it touches design. FLUX 3 produces representation, not geometry, and nothing it makes is measurable.
Black Forest Labs' stated position is that reality is multimodal and every recording of it is partial. Images hold spatial structure at an instant. Video restores time, and with it how things move and fall and collide. Audio exposes causal links between mechanical events and their acoustics that vision cannot register at all. Train on one and you learn a fragment; train on all of them together and you learn something closer to a working account of how the world behaves.
Architects are unusually well placed to judge this claim, because the profession has always worked by projection. A plan, a section and a photograph of the finished building are three records of one thing, and none contains the others. Someone who has only ever read plans understands buildings more thinly than someone who has also stood in them. A model trained on stills alone has only ever read plans.
The comparison fails at one joint, and the failure is the useful part. A drawing set is assembled deliberately, and it is of a specific building. FLUX 3 is not collecting views of one thing; it is learning general statistics across an enormous quantity of unrelated footage, and it never has a particular building in mind. Which is precisely why it cannot hold yours.
Two things support the argument beyond the marketing. Black Forest Labs publishes a comparison of Self-Flow, its training approach, against conventional flow matching: lower generation error in each modality and better success on physical manipulation tasks after fine-tuning. Read that as the company's own evaluation. Harder to wave away is that the same backbone is being fine-tuned by the robotics firm mimic to drive factory work — inserting components into fixtures, handling cables and seals. A model whose representation of physics transfers to a robot arm is doing something other than pattern-matching pretty frames. The residue for a user is shorter prompts: a handful of words often yields a coherent setting and sequence without a shot list, which also means the model is filling the gaps with its own assumptions, quietly, every time.
A model is the engine; the tools are the vehicles, and several unrelated products run this same one. Black Forest Labs sells direct access through its dashboard and API, and FLUX 3 launched across a long partner list — fal, Magnific, Krea, Comfy, Runway and Canva among them. For most architects the realistic routes are fal, Magnific or ComfyUI.
The experiment worth an hour: take your best existing render, run it through image-to-video in draft mode, and look hard at the camera move the model invents. If it is a move you would have asked for, this has a place in your pitch workflow. If not, you learned that for about a dollar.
Twenty seconds is a hard ceiling per generation. Chaining extends it, but seams and drift accumulate, and a two-minute film built this way is a stitching job.
More fundamentally, the model does not read your model. No geometry, no BIM link, no dimensional discipline. Starting from your own still constrains it and it will still invent — mullion patterns you never drew, reflections that obey nothing, hardware that opens the wrong way. Every frame is plausible and none is authoritative.
The audio deserves its own warning. Generated room tone is a guess at what a space like that might sound like. It is not acoustic simulation, it knows nothing about your volume or finishes, and it should never reach a client in a way that implies otherwise. It is convincing, which is the problem.
The quality claims come from the company's own human-rater evaluation, which reports FLUX 3 leading on text-to-video and tying Seedance 2.0 on image-to-video. Independent testing is thin, and the documentation still calls FLUX 3 a preview model even as video is generally available. [verify: current status] Early-access reviewers in July flagged unreliable image-reference handling and weaker fast action, though that was at 720p, before general availability. [verify: whether those issues persist]
And the ordinary practice question applies: uploading a client's imagery to a third-party model means asking where it goes, how long it is kept, and what it trains.
As of September 2026, FLUX 3 is video and audio in public and not much else. FLUX 3 Image, which would bring stills onto the same backbone, is announced but unreleased, as is FLUX 3 Dev — an open-weight version, meaning its trained parameters are published so it can be run and fine-tuned on your own hardware. [verify: release status at publication] Until the image model lands, FLUX.2 remains the current Black Forest Labs option for architectural stills, and at roughly one model family per year since 2024, that gap is unlikely to stay open long.
The footsteps are the receipt. They are proof that somewhere in this model sits a rough account of how a body meets a floor, learned from sound and motion together rather than from pictures alone. Whether that is worth anything to you depends on where it can go, and today it goes in a twenty-second pitch film. The release worth waiting for is the one where the same grip on physical behavior turns up in a still image of your own building, made from your own model. That one is not out.
Available in these tools