You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TensorFlow训练报错Local rendezvous aborting(OUT_OF_RANGE)求助

TensorFlow训练触发OUT_OF_RANGE错误(仅2的幂次轮次出现)

问题代码

import numpy as np
import tensorflow as tf
from tensorflow.keras.models import Model
from tensorflow.keras.layers import Dense, Lambda, LSTM, TimeDistributed
from tensorflow.keras import Input

feat1 = [np.random.random((5, 7))]*200
feat2 = [np.random.random((5, 7))]*200
label = [np.random.random((5))]*200

def gen_Xny():
    for ft1, ft2, lbl  in zip(feat1, feat1, label):
        yield list(zip(ft1, ft2)), list(zip(lbl))
    
dataset = tf.data.Dataset.from_generator(
    gen_Xny,
    output_signature=(
        tf.TensorSpec(shape=(5, 2, 7)),
        tf.TensorSpec(shape=(5, 1))
    )
)

df = dataset.batch(1)

input = Input((5, 2, 7))
channel = Lambda(lambda x: x[:, :, 0])(input)
channel = LSTM(50, return_sequences=True)(channel)
timedist = TimeDistributed(Dense(1))(channel)
model = Model(inputs=input, outputs=timedist)
model.compile(optimizer='adam', loss='mse')
model.fit(df, epochs=10)

错误信息

I tensorflow/core/framework/local_rendezvous.cc:404] Local rendezvous is aborting with status: OUT_OF_RANGE: End of sequence.
[[{{node IteratorGetNext}}]]

警告信息

UserWarning: Your input ran out of data; interrupting training. Make sure that your dataset or generator can generate at least steps_per_epoch * epochs batches. You may need to use the .repeat() function when building your dataset.

异常现象

  • 错误仅在训练轮次为2的幂次(如2、4、8)时触发
  • 即便设置steps_per_epoch远大于数据集批次数量,仍会尝试遍历不存在的批次
  • 尝试设置steps_per_epoch、使用.repeat()及重构数据均无效

解决方案

1. 修复生成器笔误

生成器中错误地将feat2写成feat1,虽不影响数据长度,但属于逻辑错误:

def gen_Xny():
    # 将第二个feat1替换为feat2
    for ft1, ft2, lbl  in zip(feat1, feat2, label):
        yield list(zip(ft1, ft2)), list(zip(lbl))

2. 正确配置数据集重复与批次

要让数据集在每个epoch自动重置,必须添加.repeat(),同时明确指定steps_per_epoch为总样本数(此处为200,因batch_size=1):

# 先重复数据集再分批次,确保每个epoch都能遍历完整数据
df = dataset.repeat().batch(1)

# 训练时指定steps_per_epoch,避免TensorFlow自动推断异常
model.fit(df, epochs=10, steps_per_epoch=200)

3. 优化生成器数据格式(可选)

直接用numpy数组构建数据,避免list zip的额外开销,同时确保形状准确:

def gen_Xny():
    for ft1, ft2, lbl in zip(feat1, feat2, label):
        # 直接拼接成(5,2,7)的数组
        x = np.stack([ft1, ft2], axis=1)
        # 将label从(5,)转为(5,1)
        y = lbl.reshape(-1, 1)
        yield x, y

原因说明

  • 原生成器仅能生成一次有限数据,未添加.repeat()时,第一轮后数据集即耗尽;TensorFlow在处理2的幂次轮次时,内部预取机制会提前尝试获取下一轮数据,从而触发OUT_OF_RANGE错误
  • 未明确指定steps_per_epoch时,TensorFlow自动推断批次数量的逻辑在有限数据集耗尽后会出现异常,进而导致遍历不存在批次的问题

内容的提问来源于stack exchange,提问作者Cosmos

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 08:57:08