FLUX 3 is now in Early Access! Exciting new multimodal learning ahead.
FLUX 3 - Real World Models: Towards Multimodal Flow Models
Original: FLUX 3 - Real World Models: Towards Multimodal Flow Models as the Backbone of Visual Intelligence.
Importance: 新しいマルチモーダル基盤モデルの発表は多くのユーザーに影響するため。
Summary
FLUX 3 is a new multimodal foundation model that integrates learning from images, videos, and audio. It aims to build a representation of the world that single modalities cannot capture, enabling perception, prediction, and action across physical and digital environments. It is currently available in Early Access, with promising early results in content creation and physical AI.
Key Points
- FLUX 3 integrates learning from images, audio, and video
- Early results are promising in content creation
- Strengthened multilingual capabilities
- Available in Early Access
- Strong in capturing human facial expressions
View developer notes (APIs, breaking changes, migration)
FLUX 3 is a new multimodal foundation model that learns from images, videos, and audio simultaneously, based on the Self-Flow approach. This has enabled a significant scaling of computational resources for training across modalities. FLUX 3 can generate outputs from text prompts and reference images/videos, showing high preference in early evaluations compared to other models, particularly excelling in capturing human facial expressions, associating sounds with physical events, and multilingual capabilities.
Source: https://bfl.ai/blog/flux-3
Outlet: Black Forest Labs
This article is an AI-generated summary (OpenAI GPT-4o-mini) of publicly available information from Anthropic, OpenAI, Google, Meta, Mistral, DeepSeek, Sakana, and other vendors. The original source URL is always provided in accordance with fair-use citation requirements. Summaries are AI-generated and may contain mistranslations or misinterpretations. Always verify details with the original source.