CROSSBFM Freezes Source Latent Space to Distill Multi-Robot Behavior in One Hour
Tech Walker reported on October 9 about a paper that proposes CROSSBFM, a method for cross-robot behavior transfer. Behavior foundation models can use a single latent variable vector to represent diverse objectives, such as imitating a motion, reaching a target pose, or maximizing a reward.
Training one robot, however, consumes hundreds of GPU hours. More troublesome is that when a second robot is trained, its new latent space is unrelated to the first: the same coordinate point corresponds to completely different actions. This mismatch makes direct transfer impossible.
CROSSBFM addresses this by freezing the source robot's behavior foundation model latent space and using a unified encoder to treat retargeted data as cross-robot correspondences. In effect, it builds a bridge between two coordinate systems without re-binarizing the entire book. The result is that multi-robot behavior spaces can be distilled within one hour.
The paper examines why latent spaces cannot be transferred directly, how the unified encoder is designed, how policy training and reward optimization are implemented across robots, and where the boundaries of generalization lie.
Retargeted data, originally generated by mapping motion between embodiments, become positive pairs that align the two latent spaces. The encoder learns a mapping that preserves task-relevant structure while filtering embodiment-specific noise. Once aligned, policies trained in the source space can be deployed on new robots with limited fine-tuning.
Experiments show one-hour distillation, but the article also discusses limitations: alignment depends on the quality of retargeted correspondences; highly different morphologies may still degrade performance; and reward optimization may require adaptation. By reusing a frozen source space, the method reduces the need to retrain each new robot from scratch and points toward more scalable behavior foundation models.