You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何自定义input_features后LeRobot策略仍忽略额外摄像头流?

问题描述

我正在用LeRobot训练SO101机械臂的策略,用到3个视频流(front、above、gripper)和一个状态向量,对应的Hugging Face数据集包含所有这三个流。我创建了自定义的train_config.json,在policy.input_features里明确列出了三个视觉流,甚至设置了"use_policy_training_preset": false禁用预设配置加载,但训练时策略只考虑front观测,完全忽略另外两个流。之前的多流训练黑客松项目用预设配置就能正常运行,所以禁用预设不是必须的。

我通过--config_path参数把自定义配置传入lerobot.scripts.train,训练初始打印的配置里包含所有三个流,但训练完成后,上传到Hugging Face的模型仓库里的train_config.json只保留了以下内容:

"input_features": {
    "observation.state": { ... },
    "observation.images.front": { ... },
"output_features": { ... }

above和gripper流丢失了,明明数据集里有这些流,我也在自定义配置里明确配置了,想问下是哪个内部步骤或配置覆盖了自定义的input_features?怎么确保LeRobot用所有提供的视频流训练?

我的自定义train_config.json内容如下:

{
    "dataset": {
        "repo_id": "aaron-ser/SO101-Dataset",
        "root": null,
        "episodes": null,
        "image_transforms": {
            "enable": false,
            "max_num_transforms": 3,
            "random_order": false,
            "tfs": {
                "brightness": {
                    "weight": 1.0,
                    "type": "ColorJitter",
                    "kwargs": {
                        "brightness": [
                            0.8,
                            1.2
                        ]
                    }
                },
                "contrast": {
                    "weight": 1.0,
                    "type": "ColorJitter",
                    "kwargs": {
                        "contrast": [
                            0.8,
                            1.2
                        ]
                    }
                },
                "saturation": {
                    "weight": 1.0,
                    "type": "ColorJitter",
                    "kwargs": {
                        "saturation": [
                            0.5,
                            1.5
                        ]
                    }
                },
                "hue": {
                    "weight": 1.0,
                    "type": "ColorJitter",
                    "kwargs": {
                        "hue": [
                            -0.05,
                            0.05
                        ]
                    }
                },
                "sharpness": {
                    "weight": 1.0,
                    "type": "SharpnessJitter",
                    "kwargs": {
                        "sharpness": [
                            0.5,
                            1.5
                        ]
                    }
                }
            }
        },
        "revision": null,
        "use_imagenet_stats": true,
        "video_backend": "torchcodec"
    },
    "env": null,
    "policy": {
        "type": "act",
        "n_obs_steps": 1,
        "normalization_mapping": {
            "VISUAL": "MEAN_STD",
            "STATE": "MEAN_STD",
            "ACTION": "MEAN_STD"
        },
        "input_features": {
            "observation.state": {
                "type": "STATE",
                "shape": [
                    6
                ]
            },
            "observation.images.front": {
                "type": "VISUAL",
                "shape": [
                    3,
                    720,
                    1280
                ]
            },
            "observation.images.above": {
                "type": "VISUAL",
                "shape": [
                    3,
                    720,
                    1280
                ]
            },
            "observation.images.gripper": {
                "type": "VISUAL",
                "shape": [
                    3,
                    720,
                    1280
                ]
            }
        },
        "output_features": {
            "action": {
                "type": "ACTION",
                "shape": [
                    6
                ]
            }
        },
        "device": "cuda",
        "use_amp": false,
        "push_to_hub": true,
        "repo_id": "aaron-ser/SO101-Model",
        "private": null,
        "tags": null,
        "license": null,
        "chunk_size": 100,
        "n_action_steps": 100,
        "vision_backbone": "resnet18",
        "pretrained_backbone_weights": "ResNet18_Weights.IMAGENET1K_V1",
        "replace_final_stride_with_dilation": false,
        "pre_norm": false,
        "dim_model": 512,
        "n_heads": 8,
        "dim_feedforward": 3200,
        "feedforward_activation": "relu",
        "n_encoder_layers": 4,
        "n_decoder_layers": 1,
        "use_vae": true,
        "latent_dim": 32,
        "n_vae_encoder_layers": 4,
        "temporal_ensemble_coeff": null,
        "dropout": 0.1,
        "kl_weight": 10.0,
        "optimizer_lr": 1e-05,
        "optimizer_weight_decay": 0.0001,
        "optimizer_lr_backbone": 1e-05
    },
    "output_dir": "/Users/aaronserpilin/Documents/Storks/storks-so101/my_robot/outputs/train/multiple_stream_test",
    "job_name": "multiple_stream_test",
    "resume": false,
    "seed": 1000,
    "num_workers": 4,
    "batch_size": 8,
    "steps": 1000,
    "eval_freq": 20000,
    "log_freq": 200,
    "save_checkpoint": true,
    "save_freq": 20000,
    "use_policy_training_preset": false,
    "optimizer": {
        "type": "adamw",
        "lr": 1e-05,
        "weight_decay": 0.0001,
        "grad_clip_norm": 10.0,
        "betas": [
            0.9,
            0.999
        ],
        "eps": 1e-08
    },
    "scheduler": {
        "type": "cosine_decay_with_warmup",
        "num_warmup_steps": 10000,
        "num_decay_steps": 90000,
        "peak_lr": 1e-4,
        "decay_lr": 2.5e-6
    },
    "eval": {
        "n_episodes": 50,
        "batch_size": 50,
        "use_async_envs": false
    },
    "wandb": {
        "enable": true,
        "disable_artifact": false,
        "project": "lerobot",
        "entity": null,
        "notes": null,
        "mode": null
    }
}
问题分析与解决方法

可能的原因

  1. 数据集元数据自动对齐:LeRobot加载数据集时会读取根目录下的dataset_info.json,自动提取里面定义的观测特征列表,如果该文件只标注了front视觉流,训练脚本会自动过滤配置里的其他流,覆盖自定义的input_features。
  2. ACT策略模型的单流限制:你使用的act类型策略默认只支持单视觉输入,模型代码里硬编码了仅处理observation.images.front,即便配置里添加了其他流,模型也不会读取。
  3. 配置保存时的特征过滤:训练结束保存配置时,脚本可能会根据模型实际接收的输入特征反向生成input_features字段,而非直接保存传入的原始配置。

解决步骤

1. 检查并更新数据集元数据

  • 打开数据集根目录下的dataset_info.json,确认observation字段里的images包含front、above、gripper三个键,格式示例:
    "observation": {
        "state": {"shape": [6], "dtype": "float32"},
        "images": {
            "front": {"shape": [3, 720, 1280], "dtype": "uint8"},
            "above": {"shape": [3, 720, 1280], "dtype": "uint8"},
            "gripper": {"shape": [3, 720, 1280], "dtype": "uint8"}
        }
    }
    
  • 若缺少对应字段,补充完整后重新上传数据集。

2. 修改ACT模型以支持多视觉流

找到LeRobot代码中的act.py模型文件(通常在lerobot/models/目录下),修改特征提取逻辑:

  • 将原来仅提取front图像特征的代码,改为遍历所有视觉流:
    # 替换原单流特征提取代码
    visual_features = []
    for cam_name in self.visual_feature_names:
        img = observations[f"observation.images.{cam_name}"]
        feat = self.vision_backbone(img)
        visual_features.append(feat)
    # 拼接所有视觉特征
    visual_features = torch.cat(visual_features, dim=-1)
    
  • 同时在模型初始化阶段,从配置的input_features中读取所有视觉流名称,而非硬编码front。

3. 禁用配置的自动对齐逻辑

在训练脚本train.py中,找到加载数据集后更新配置的代码段(通常类似config.policy.input_features = dataset.features),注释掉该逻辑,确保自定义配置不会被数据集特征覆盖。

4. 修正配置保存逻辑

检查保存配置的代码(通常在train.py结尾或lerobot/utils/config.py中),确保直接保存传入的原始config对象,而非从模型或数据集生成的简化版本。

内容的提问来源于stack exchange,提问作者Aaron Serpilin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 14:42:32