You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Ray Tune调参时BroadModel出现AttributeError问题求助

问题:Ray Tune中扩大网络规模或增加迭代数出现AttributeError: 'BroadModel' object has no attribute 'model'

问题场景

自定义BroadModel继承tune.Trainable,使用Population Based Training(PBT)进行超参数调优。当网络各层规模扩大到100及以上,或训练迭代数设置为40以上时,触发如下错误:

Failure # 1 (occurred at 2022-09-05_12-04-07)
[36mray::ResourceTrainable.train()[39m (pid=35719, ip=192.168.91.120, repr=<ray.tune.trainable.util.BroadModel object at 0x7f478f107c40>)
File "/home/ssrc/asq/lib/python3.8/site-packages/ray/tune/trainable/trainable.py", line 347, in train
result = self.step()
File "ray_test.py", line 258, in step
self.model.fit( AttributeError: 'BroadModel' object has no attribute 'model'

而当迭代数设置在20以内、网络规模较小时,训练正常运行。

错误原因分析

  1. 模型初始化超时:网络规模扩大后,模型构建、编译的时间变长,Ray Actor可能因超时未完成setup方法就进入step,导致self.model未被赋值。
  2. 全局变量冲突:build_model中使用了全局变量(convB2、drop2等),在PBT复用Actor或多trial并行时,会导致变量状态混乱,可能中断模型初始化流程。
  3. 数据加载冗余:每次调用build_model都重新加载训练/测试数据,不仅浪费资源,还可能因内存占用过高导致Actor初始化失败。
  4. setup方法返回值问题:setup方法返回了model,但Ray Trainable的setup不需要返回值,多余的返回可能干扰内部逻辑。

解决方案

1. 重构数据加载逻辑

将数据加载移到setup方法中,仅执行一次,避免重复加载和内存浪费:

class BroadModel(tune.Trainable):
    os.environ['TF_CPP_MIN_LOG_LEVEL'] = '3'
    
    def setup(self, config):
        # 只加载一次数据,避免重复操作
        window_size = 200
        self.x_gyro, self.x_acc, x_mag, q = load_data_train()
        self.Att_quat = Att_q(q)
        self.x_gyro_t, self.x_acc_t, x_mag_t, q_t = load_data_test()
        self.Att_quat_t = Att_q(q_t)
        self.x_gyro, self.x_acc, self.Att_quat = shuffle(self.x_gyro, self.x_acc, self.Att_quat)
        
        model = self.build_model(config, window_size)
        model.compile(
            optimizer=Adam(learning_rate=config['lr']),
            loss=quaternion_mean_multiplicative_error,
            metrics=[quaternion_mean_multiplicative_error],
        )
        self.model = model
        # 移除多余的return语句

2. 移除全局变量

删除build_model中的全局变量声明,改用局部变量或类属性,避免多Actor冲突:

def build_model(self, config, window_size):
    # 移除global声明
    x1 = Input((window_size, 3), name='x1')
    x2 = Input((window_size, 3), name='x2')
    
    convA1 = Conv1D(config["Conv1DA"], 11, padding='same', activation='relu')(x1)
    # 优化循环逻辑,避免i>0的冗余判断
    current_layer = convA1
    for i in range(1, config["Conv1DAn"] + 1):
        current_layer = Conv1D(config[f'Conv1DAn_{i}'], 11, padding='same', activation='relu')(current_layer)
    poolA = MaxPooling1D(3)(current_layer)
    
    convB1 = Conv1D(config["Conv1DB"], 11, padding='same', activation='relu')(x2)
    current_layer = convB1
    for i in range(1, config["Conv1DBn"] + 1):
        current_layer = Conv1D(config[f'Conv1DBn_{i}'], 11, padding='same', activation='relu')(current_layer)
    poolB = MaxPooling1D(3)(current_layer)
    
    AB = concatenate([poolA, poolB])
    
    lstm1 = Bidirectional(LSTM(config["LSTM1"], return_sequences=True))(AB)
    drop1 = Dropout(config['dropout'])(lstm1)
    for i in range(1, config['LSTMn'] + 1):
        lstm2 = Bidirectional(LSTM(config[f'LSTMn_{i}'], return_sequences=True))(drop1)
        drop1 = Dropout(config['dropout'])(lstm2)   
    lstm2 = Bidirectional(LSTM(config['LSTMn_l']))(drop1)
    drop2 = Dropout(config['dropout'])(lstm2)
    y1_pred = Dense(4, kernel_regularizer='l2')(drop2)
    model = Model(inputs=[x1, x2], outputs=[y1_pred])
    return model

3. 调整Ray Actor超时设置

在初始化Ray时增加超时配置,给模型初始化足够时间:

if __name__ == "__main__":
    import ray
    ray.init(
        runtime_env={"env_vars": {"RAY_TRAINABLE_SETUP_TIMEOUT": "300"}},  # 设置5分钟超时
        ignore_reinit_error=True
    )
    # 后续PBT、Tuner代码不变

4. 确保step方法安全访问self.model

在step方法中增加判断,避免未初始化时调用self.model.fit:

def step(self):
    if not hasattr(self, 'model'):
        # 模型未初始化,返回错误结果或重新初始化
        return {"loss": float("inf"), "training_iteration": self.iteration}
    
    # 原有训练逻辑
    history = self.model.fit(
        [self.x_gyro, self.x_acc],
        self.Att_quat,
        batch_size=self.config['batch_size'],
        epochs=self.config['epochs'],
        validation_data=([self.x_gyro_t, self.x_acc_t], self.Att_quat_t),
        verbose=0
    )
    return {"loss": history.history['loss'][-1], "val_loss": history.history['val_loss'][-1], "training_iteration": self.iteration}

5. 优化资源配置

根据模型规模调整resources_per_trial,确保每个trial有足够的CPU内存:

resources_per_trial = {"cpu": 10, "gpu": 0, "memory": 32 * 1024}  # 增加内存配额,单位MB

修改后完整BroadModel示例

class BroadModel(tune.Trainable):
    os.environ['TF_CPP_MIN_LOG_LEVEL'] = '3'
    
    def setup(self, config):
        window_size = 200
        # 加载数据
        self.x_gyro, self.x_acc, x_mag, q = load_data_train()
        self.Att_quat = Att_q(q)
        self.x_gyro_t, self.x_acc_t, x_mag_t, q_t = load_data_test()
        self.Att_quat_t = Att_q(q_t)
        self.x_gyro, self.x_acc, self.Att_quat = shuffle(self.x_gyro, self.x_acc, self.Att_quat)
        
        # 构建并编译模型
        model = self.build_model(config, window_size)
        model.compile(
            optimizer=Adam(learning_rate=config['lr']),
            loss=quaternion_mean_multiplicative_error,
            metrics=[quaternion_mean_multiplicative_error],
        )
        self.model = model
    
    def build_model(self, config, window_size):
        x1 = Input((window_size, 3), name='x1')
        x2 = Input((window_size, 3), name='x2')
        
        # 构建ConvA分支
        convA1 = Conv1D(config["Conv1DA"], 11, padding='same', activation='relu')(x1)
        current_layer = convA1
        for i in range(1, config["Conv1DAn"] + 1):
            current_layer = Conv1D(config[f'Conv1DAn_{i}'], 11, padding='same', activation='relu')(current_layer)
        poolA = MaxPooling1D(3)(current_layer)
        
        # 构建ConvB分支
        convB1 = Conv1D(config["Conv1DB"], 11, padding='same', activation='relu')(x2)
        current_layer = convB1
        for i in range(1, config["Conv1DBn"] + 1):
            current_layer = Conv1D(config[f'Conv1DBn_{i}'], 11, padding='same', activation='relu')(current_layer)
        poolB = MaxPooling1D(3)(current_layer)
        
        AB = concatenate([poolA, poolB])
        
        # 构建LSTM分支
        lstm1 = Bidirectional(LSTM(config["LSTM1"], return_sequences=True))(AB)
        drop1 = Dropout(config['dropout'])(lstm1)
        for i in range(1, config['LSTMn'] + 1):
            lstm2 = Bidirectional(LSTM(config[f'LSTMn_{i}'], return_sequences=True))(drop1)
            drop1 = Dropout(config['dropout'])(lstm2)   
        lstm2 = Bidirectional(LSTM(config['LSTMn_l']))(drop1)
        drop2 = Dropout(config['dropout'])(lstm2)
        y1_pred = Dense(4, kernel_regularizer='l2')(drop2)
        
        model = Model(inputs=[x1, x2], outputs=[y1_pred])
        return model
    
    def step(self):
        if not hasattr(self, 'model'):
            return {"loss": float("inf"), "training_iteration": self.iteration}
        
        history = self.model.fit(
            [self.x_gyro, self.x_acc],
            self.Att_quat,
            batch_size=self.config['batch_size'],
            epochs=self.config['epochs'],
            validation_data=([self.x_gyro_t, self.x_acc_t], self.Att_quat_t),
            verbose=0
        )
        return {
            "loss": history.history['loss'][-1],
            "val_loss": history.history['val_loss'][-1],
            "training_iteration": self.iteration
        }

内容的提问来源于stack exchange,提问作者Arman Asgharpoor

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 21:05:22