Revolutionary Flux 3 Model Generates 20-Second Videos with Native Audio, Leaving Rivals in the Dust
Black Forest Labs' latest multimodal foundation model, Flux 3, has achieved a groundbreaking milestone by generating videos up to 20 seconds long with native audio, outperforming several rival models in early tests. This innovation has significant implications for developers, businesses, and everyday users, marking a major step towards real-world visual intelligence.
The German AI company Black Forest Labs has made a significant breakthrough in the field of artificial intelligence with the release of its Flux 3 model, a multimodal foundation model that learns from images, videos, and audio simultaneously. This model has achieved a major milestone by generating videos up to 20 seconds long with native audio, a first for the company. The Flux 3 model supports a range of features, including text-to-video, image-to-video, video-to-video, keyframe-based transitions, multilingual dialogue, and agent-driven links between clips for longer multi-shot sequences. In early evaluations, the model has outperformed several rival models, including Luma Ray 3.2, Runway Gen-4.5, and Grok Imagine Video, with user preference rates of 93%, 77%, and 69%, respectively.
The Flux 3 model has also been compared to stronger competitors, including Kling v3 Pro, Happy Horse v1, and Seedance 2.0, with user preference rates of 60%, 59%, and 52%, respectively. While the margins are narrower against these competitors, the results are still impressive, especially considering that the model is still in its early stages. The Flux 3 model is expected to improve image generation, especially for complex prompts and accurate text rendering, making it a powerful tool for developers and businesses. The model's ability to generate high-quality videos with native audio has significant implications for a range of applications, including advertising, education, and entertainment.
The release of the Flux 3 model is part of a broader push to build so-called world models, which are designed to perceive, predict, and act across physical and digital environments. The company argues that no single modality can capture reality in full, and that training on multiple modalities simultaneously allows the model to fill gaps and gain a more comprehensive understanding of the world. The Flux 3 model is a major step towards achieving this goal, and its release has significant implications for the field of artificial intelligence. The model's ability to generate high-quality videos with native audio has the potential to revolutionize a range of applications, from advertising and education to entertainment and beyond.
In addition to the Flux 3 model, Black Forest Labs has also developed Flux-mimic, a video action model for robotics applications, which is already being tested at Audi. This model has the potential to revolutionize the field of robotics, enabling robots to perceive and interact with their environment in a more sophisticated way. The release of the Flux 3 model and Flux-mimic demonstrates the company's commitment to pushing the boundaries of artificial intelligence and developing innovative solutions that have the potential to transform a range of industries.
The release of the Flux 3 model has significant implications for AI model users and developers, who will be able to leverage the model's capabilities to generate high-quality videos with native audio. This has the potential to open up new opportunities for applications such as video generation, robotics, and beyond. As the field of artificial intelligence continues to evolve, the release of the Flux 3 model is a major milestone, demonstrating the potential for multimodal foundation models to achieve real-world visual intelligence. With its impressive performance and range of features, the Flux 3 model is set to revolutionize the field of artificial intelligence, and its impact will be felt for years to come.