模型

RT-2

Google DeepMind 视觉-语言-动作模型,将 VLM 与机器人控制统一为 token 预测。

模型

主体信息

Google DeepMind 视觉-语言-动作模型,将 VLM 与机器人控制统一为 token 预测。

相关论文