Studio product
Interface preview
Create a video directly from a detailed text prompt.
Your prompt is sent to the video generation service when you submit.
Interface preview
Interface preview
Interface preview
Interface preview
Create video from words, images, start and end frames, or a reference clip—with motion, camera direction and native audio.
FLUX 3 Video accepts text, images, reference video and ordered keyframes. It can generate multi-shot clips with optional speech, effects and ambience, and it can continue an existing clip.
Video capability notes are drawn from Black Forest Labs announcements. Shot-planning advice reflects practical production workflows; showcase clips are illustrative and are not presented as controlled quality benchmarks.
Describe the subject, visible action, environment, camera behavior, style and sound in one focused brief.
Begin from a visual reference and direct motion while preserving the identity, composition and art direction that matter.
Supply a starting image and ending image to guide composition and the transition through the shot.
Use a source clip to guide central elements, or extend an existing video with matching continuity.
Explore product films, campaign concepts, motion design and social creative before committing a full production.
Turn scripts and visual references into moving previs with camera language, performance and pacing.
Develop cinematic, documentary, anime, VFX, nature and experimental visual worlds.
Choose text, an image, ordered frames or a reference clip according to the information the shot must preserve.
Capability claims are linked to their source; illustrative clips demonstrate workflow ideas without being presented as controlled model tests.
Review motion continuity, camera logic, synchronized sound and the ending frame across the complete sequence before delivery.
Define the subject, action, location, framing and closing moment first. Then pick the right input—text, image, keyframes or reference clip—based on what the shot must preserve.
Describe the subject move before the camera move. Add environment, light, pace and sound; keep dialogue brief and pair each sound cue with a visible event.
Use a reference for identity, a clip for timing, and start/end frames for the transition. Give each asset a single role—fewer purposeful anchors beat a pile of overlapping references.
Play the whole output to check identity drift, geometry warping, camera jumps, unstable text and sync. Approve by playback, then inspect the transition into the final frame.
When a clip is close, keep the subject and scene while changing only one thing—camera path, action speed, transition or ending. Record the prompt, inputs, duration and audio settings for repeatable results.
FLUX 3 Video is the video-generation part of Black Forest Labs’ multimodal FLUX 3 system. Published capabilities include video from text, images, keyframes and reference inputs, continuation from existing video, and optional native audio.
Choose Text to Video for a written scene, Image to Video when an approved still should anchor the shot, or Reference to Video when an input should guide motion or continuity. Enter the required prompt and source, select duration and resolution, review the credit estimate, and submit.
The model interprets the prompt together with any supplied visual context as a timed sequence. Instructions can define subject action, environment, camera behavior, lighting, pace, dialogue, effects, ambience, and the intended final moment.
Start with the subject and visible action, then add environment, framing, camera movement, lighting, pace, sound, and the ending. Use observable instructions, keep dialogue concise, and change one major variable at a time when refining a result.
Yes. Text-to-video begins with a written scene brief and does not require a source image. A useful prompt describes what happens, where it happens, how the camera moves, and what should be heard.
Yes. Image-to-video uses a still image as the visual starting point. The image can anchor identity, composition, product appearance, wardrobe, or style while the prompt directs motion, camera behavior, sound, and the ending.
Published FLUX 3 capabilities include video continuation. The reference workflow can extend an existing clip, but the transition still needs review for identity, geometry, movement direction, lighting, ambience, and pacing.
Black Forest Labs states that FLUX 3 can generate video with optional native audio for up to 20 seconds in one generation. Audio can include multilingual dialogue, synchronized effects, and environmental ambience; provider settings may vary.
Choose a video workflow, confirm the settings and review the credit estimate before submitting.
Open the generator →