ATRIUMsearch → argument graph
Article · 2026-07-24 · 6 moments

[AINews] Black Forest Labs FLUX 3 - Multimodal Flow Models that beat Seedance 2.0, Gemini Omni and Grok Imagine, and FLUX-mimic video-action robotics model

A HUGE win for BFL! ✦ AI generated

01
Claim

Better video world modeling transfers directly into robot control quality and sample efficiency, as demonstrated by FLUX-mimic's concrete robotics instantiation on a single on-prem GPU.

mimic's FLUX-mimic is a concrete robotics instantiation built on FLUX 3, trained on robot and wearable data for general-purpose dexterity and deployable on a single on-prem GPU. The central claim is that better video world modeling directly transfers to robot control quality and sample efficiency, with testing already underway at Audi.

transcript

AINews / Latent.Space: mimic's FLUX-mimic is a concrete robotics instantiation of that thesis: @mimicrobotics described FLUX-mimic as a Video-Action Model built on top of FLUX 3, trained on robot and wearable data for general-purpose dexterity and deployable on a single on-prem GPU. Their central claim is that better video world modeling transfers directly into robot control quality and sample efficiency; they're already testing with Audi.

gives example · 1

02
Claim

FLUX 3 is a unified multimodal model spanning image, video, audio, and action prediction, with an architecture intended to bridge media generation and control — not a loose family of specialized generators.

BFL's FLUX 3 is a unified multimodal model covering image, video, audio, and action prediction, with team members connecting it back to Self-Flow research. The unified training story is that one architecture is intended to bridge media generation and control, not just a collection of specialized generators.

transcript

AINews / Latent.Space: Black Forest Labs' FLUX 3 expands the multimodal frontier beyond image/video: @bfl_ai launched FLUX 3, a unified multimodal model spanning image, video, audio, and action prediction, with early access for FLUX 3 Video and an explicit claim that the same architecture can be extended toward robotics. Team members connected it back to the earlier Self-Flow research, including @hila_chefer and @robrombach. What matters technically is the unified training story: not a loose family of specialized generators, but one architecture intended to bridge media generation and control.

03
Context

Health in ChatGPT is a strategically important rollout that builds a new high-trust application layer on top of existing model capability, with additional encryption and a commitment not to use health data for training or ads.

OpenAI's Health in ChatGPT rollout allows US users to connect Apple Health and supported medical records, with additional encryption, a commitment not to train foundation models or target ads, and substantial physician review effort. The text argues this is less about a new model and more about a new high-trust application layer on top of existing model capability.

transcript

AINews / Latent.Space: Health in ChatGPT is a more strategically important rollout than it may first appear: @OpenAI, @ChatGPTapp, and @thekaransinghal announced U.S. rollout of Health in ChatGPT, allowing users to connect Apple Health and supported medical records. The notable implementation claims: connected health data receives additional encryption, is not used to train foundation models or target ads, and the feature builds on substantial physician review effort. This is less about a new model and more about a new high-trust application layer on top of existing model capability.

gives example · 1

04
Context

Open datasets like The Stack v3 materially raise the floor for every lab that wants to build competitive code models without relying on closed ecosystems, and open weights are strategically important against calls to restrict distillation.

Multiple high-signal posts pushed back on attempts to sharply separate internet-scale pretraining from output-level distillation. The subtext is that open datasets like The Stack v3 materially raise the floor for every lab that wants to build competitive code models without relying on closed ecosystems, and open weights are strategically important.

transcript

AINews / Latent.Space: Distillation remains the live ideological fault line: several high-signal posts pushed back on attempts to sharply separate 'internet-scale pretraining' from output-level distillation. @GergelyOrosz compared model inspection via prompting to reverse-engineering a competitor's product, while @SchmidhuberAI emphasized distillation's long lineage. @Suhail argued the practical response is not prohibition but stronger investment in open-weight domestic models, and @garrytan put it more simply: open weights are strategically important. The subtext across these posts is that open datasets like The Stack v3 materially raise the floor for every lab that wants to build competitive code models without relying on closed ecosystems.

supports · 1

05
Claim

The center of gravity in AI is shifting from prompts to harnesses: graph engineering is mostly old software architecture renamed, and most agents still do not need complex graphs unless workflows branch, verify, or require human approvals.

Multiple tweets converged on the same engineering thesis that the center of gravity is shifting from prompts to harnesses. The Turing Post made the cleaner systems point that graph engineering is mostly old software architecture renamed, and most agents still do not need complex graphs unless workflows branch, verify, or require human approvals.

transcript

AINews / Latent.Space: The center of gravity is shifting from prompts to harnesses: multiple tweets converged on the same engineering thesis. @unclebobmartin described an 'extreme constraints' workflow where trust comes from tests, QA, mutation testing, and metrics, not manual code review. @ThePrimeagen said he has become materially more positive on AI coding workflows, especially for large structural refactors. @TheTuringPost made the cleaner systems point: 'graph engineering' is mostly old software architecture renamed, and most agents still do not need complex graphs unless workflows branch, verify, or require human approvals.

06
Data

The Stack v3 is the largest open code dataset publicly released, with 114 TB raw, 224M repositories, 44B files, 770 languages, and roughly 5T deduplicated/filtered tokens.

Anton Lozhkov announced The Stack v3, now the largest open code dataset publicly released, with significant gains over v2 including a jump from ~550B to ~5T filtered tokens and large per-language increases in C++, TypeScript, Rust, and Python.

transcript

AINews / Latent.Space: @anton_lozhkov announced The Stack v3, now the largest open code dataset publicly released: 114 TB raw, 224M repositories, 44B files, 770 languages, and roughly 5T deduplicated/filtered tokens. Relative to v2, the filtered corpus jumps from ~550B to ~5T tokens, with especially large gains in C++ (x15), TypeScript (x7.5), Rust (x7), and Python (x4.8).

Highlight slides
Related episodes