Really interesting. Upshot: a well trained multimodal video generation model has a world representation model trained inside it. They’ve done some work lifting this world model out and deploying it to robots, where it seems to work well.
On the one hand, this isn’t a new idea, and the quality video models certainly have understanding of materials, light, the world (at least in an Occam’s razor sense of understanding). I’m not aware of a video lab that’s turned itself into a robot lab yet, though; perhaps this would be a first, or a new sort of obvious-in-retrospect business path: train video model, sell video generation, scale, use scale to train robot things: profit.
I found their hands very interesting - looks like a bunch of stuff hidden in gloves - Xiami’s Robotics-1 foundation model just released uses training on some pretty standard looking grippers; to the point that there are demo videos of people putting on gripper type gloves to make video to train that model.
The BFL model looks like it doesn’t need that at all. Given the difficulty of the hardware side, I’ll be curious to see what they do with this.
The video at around 3.30min, where the robot arm took 3 attempts to reseat the window trim, was quite unnerving - I have not seen such resolving before. Is it new or am I way out of the loop?
Really interesting. Upshot: a well trained multimodal video generation model has a world representation model trained inside it. They’ve done some work lifting this world model out and deploying it to robots, where it seems to work well.
On the one hand, this isn’t a new idea, and the quality video models certainly have understanding of materials, light, the world (at least in an Occam’s razor sense of understanding). I’m not aware of a video lab that’s turned itself into a robot lab yet, though; perhaps this would be a first, or a new sort of obvious-in-retrospect business path: train video model, sell video generation, scale, use scale to train robot things: profit.
I found their hands very interesting - looks like a bunch of stuff hidden in gloves - Xiami’s Robotics-1 foundation model just released uses training on some pretty standard looking grippers; to the point that there are demo videos of people putting on gripper type gloves to make video to train that model.
The BFL model looks like it doesn’t need that at all. Given the difficulty of the hardware side, I’ll be curious to see what they do with this.
The video at around 3.30min, where the robot arm took 3 attempts to reseat the window trim, was quite unnerving - I have not seen such resolving before. Is it new or am I way out of the loop?
It's nice to see partnerships between European startups.
Wasn’t Flux purchased by Meta?
No. Just a partnership: https://www.bloomberg.com/news/articles/2025-09-09/meta-to-p...
If Amazon does not buy them, the CEO should be fired...