为何自定义input_features后LeRobot策略仍忽略额外摄像头流?
我正在用LeRobot训练SO101机械臂的策略,用到3个视频流(front、above、gripper)和一个状态向量,对应的Hugging Face数据集包含所有这三个流。我创建了自定义的train_config.json,在policy.input_features里明确列出了三个视觉流,甚至设置了"use_policy_training_preset": false禁用预设配置加载,但训练时策略只考虑front观测,完全忽略另外两个流。之前的多流训练黑客松项目用预设配置就能正常运行,所以禁用预设不是必须的。
我通过--config_path参数把自定义配置传入lerobot.scripts.train,训练初始打印的配置里包含所有三个流,但训练完成后,上传到Hugging Face的模型仓库里的train_config.json只保留了以下内容:
"input_features": { "observation.state": { ... }, "observation.images.front": { ... }, "output_features": { ... }
above和gripper流丢失了,明明数据集里有这些流,我也在自定义配置里明确配置了,想问下是哪个内部步骤或配置覆盖了自定义的input_features?怎么确保LeRobot用所有提供的视频流训练?
我的自定义train_config.json内容如下:
{ "dataset": { "repo_id": "aaron-ser/SO101-Dataset", "root": null, "episodes": null, "image_transforms": { "enable": false, "max_num_transforms": 3, "random_order": false, "tfs": { "brightness": { "weight": 1.0, "type": "ColorJitter", "kwargs": { "brightness": [ 0.8, 1.2 ] } }, "contrast": { "weight": 1.0, "type": "ColorJitter", "kwargs": { "contrast": [ 0.8, 1.2 ] } }, "saturation": { "weight": 1.0, "type": "ColorJitter", "kwargs": { "saturation": [ 0.5, 1.5 ] } }, "hue": { "weight": 1.0, "type": "ColorJitter", "kwargs": { "hue": [ -0.05, 0.05 ] } }, "sharpness": { "weight": 1.0, "type": "SharpnessJitter", "kwargs": { "sharpness": [ 0.5, 1.5 ] } } } }, "revision": null, "use_imagenet_stats": true, "video_backend": "torchcodec" }, "env": null, "policy": { "type": "act", "n_obs_steps": 1, "normalization_mapping": { "VISUAL": "MEAN_STD", "STATE": "MEAN_STD", "ACTION": "MEAN_STD" }, "input_features": { "observation.state": { "type": "STATE", "shape": [ 6 ] }, "observation.images.front": { "type": "VISUAL", "shape": [ 3, 720, 1280 ] }, "observation.images.above": { "type": "VISUAL", "shape": [ 3, 720, 1280 ] }, "observation.images.gripper": { "type": "VISUAL", "shape": [ 3, 720, 1280 ] } }, "output_features": { "action": { "type": "ACTION", "shape": [ 6 ] } }, "device": "cuda", "use_amp": false, "push_to_hub": true, "repo_id": "aaron-ser/SO101-Model", "private": null, "tags": null, "license": null, "chunk_size": 100, "n_action_steps": 100, "vision_backbone": "resnet18", "pretrained_backbone_weights": "ResNet18_Weights.IMAGENET1K_V1", "replace_final_stride_with_dilation": false, "pre_norm": false, "dim_model": 512, "n_heads": 8, "dim_feedforward": 3200, "feedforward_activation": "relu", "n_encoder_layers": 4, "n_decoder_layers": 1, "use_vae": true, "latent_dim": 32, "n_vae_encoder_layers": 4, "temporal_ensemble_coeff": null, "dropout": 0.1, "kl_weight": 10.0, "optimizer_lr": 1e-05, "optimizer_weight_decay": 0.0001, "optimizer_lr_backbone": 1e-05 }, "output_dir": "/Users/aaronserpilin/Documents/Storks/storks-so101/my_robot/outputs/train/multiple_stream_test", "job_name": "multiple_stream_test", "resume": false, "seed": 1000, "num_workers": 4, "batch_size": 8, "steps": 1000, "eval_freq": 20000, "log_freq": 200, "save_checkpoint": true, "save_freq": 20000, "use_policy_training_preset": false, "optimizer": { "type": "adamw", "lr": 1e-05, "weight_decay": 0.0001, "grad_clip_norm": 10.0, "betas": [ 0.9, 0.999 ], "eps": 1e-08 }, "scheduler": { "type": "cosine_decay_with_warmup", "num_warmup_steps": 10000, "num_decay_steps": 90000, "peak_lr": 1e-4, "decay_lr": 2.5e-6 }, "eval": { "n_episodes": 50, "batch_size": 50, "use_async_envs": false }, "wandb": { "enable": true, "disable_artifact": false, "project": "lerobot", "entity": null, "notes": null, "mode": null } }
可能的原因
- 数据集元数据自动对齐:LeRobot加载数据集时会读取根目录下的
dataset_info.json,自动提取里面定义的观测特征列表,如果该文件只标注了front视觉流,训练脚本会自动过滤配置里的其他流,覆盖自定义的input_features。 - ACT策略模型的单流限制:你使用的
act类型策略默认只支持单视觉输入,模型代码里硬编码了仅处理observation.images.front,即便配置里添加了其他流,模型也不会读取。 - 配置保存时的特征过滤:训练结束保存配置时,脚本可能会根据模型实际接收的输入特征反向生成
input_features字段,而非直接保存传入的原始配置。
解决步骤
1. 检查并更新数据集元数据
- 打开数据集根目录下的
dataset_info.json,确认observation字段里的images包含front、above、gripper三个键,格式示例:"observation": { "state": {"shape": [6], "dtype": "float32"}, "images": { "front": {"shape": [3, 720, 1280], "dtype": "uint8"}, "above": {"shape": [3, 720, 1280], "dtype": "uint8"}, "gripper": {"shape": [3, 720, 1280], "dtype": "uint8"} } } - 若缺少对应字段,补充完整后重新上传数据集。
2. 修改ACT模型以支持多视觉流
找到LeRobot代码中的act.py模型文件(通常在lerobot/models/目录下),修改特征提取逻辑:
- 将原来仅提取
front图像特征的代码,改为遍历所有视觉流:# 替换原单流特征提取代码 visual_features = [] for cam_name in self.visual_feature_names: img = observations[f"observation.images.{cam_name}"] feat = self.vision_backbone(img) visual_features.append(feat) # 拼接所有视觉特征 visual_features = torch.cat(visual_features, dim=-1) - 同时在模型初始化阶段,从配置的
input_features中读取所有视觉流名称,而非硬编码front。
3. 禁用配置的自动对齐逻辑
在训练脚本train.py中,找到加载数据集后更新配置的代码段(通常类似config.policy.input_features = dataset.features),注释掉该逻辑,确保自定义配置不会被数据集特征覆盖。
4. 修正配置保存逻辑
检查保存配置的代码(通常在train.py结尾或lerobot/utils/config.py中),确保直接保存传入的原始config对象,而非从模型或数据集生成的简化版本。
内容的提问来源于stack exchange,提问作者Aaron Serpilin

