You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

训练RephraseGenerator模型触发IndexError: index out of range in self错误排查

训练自定义RephraseGenerator模型时触发IndexError: index out of range in self错误

问题描述

训练自定义RephraseGenerator模型时,训练函数抛出「IndexError: index out of range in self」错误,以下是相关代码、错误栈和数据样例,求排查问题。

主代码

# Extract input and target sequences from data list
input_sequences = []
target_sequences = []
BATCH_SIZE = 64
data = read_csv('gpt-j-data.csv')

for query, rephrases in data:
    input_sequences.append(query)
    target_sequences.append(rephrases)

# Load the tokenizer
tokenizer = GPT2Tokenizer.from_pretrained('gpt2')

# Tokenize the input and target sequences
input_sequences = [tokenizer.encode(sequence, add_special_tokens=True) for sequence in input_sequences]
target_sequences = [tokenizer.encode(sequence, add_special_tokens=True) for sequence in target_sequences]

# Convert the input and target sequences to tensors
input_sequences = [torch.tensor(sequence) for sequence in input_sequences]
target_sequences = [torch.tensor(sequence) for sequence in target_sequences]

input_sequences = ensure_tensor_size(input_sequences, 4)
target_sequences = ensure_tensor_size(target_sequences, 4)

# Create a RephraseDataset object from the input and target sequences
dataset = RephraseDataset(input_sequences, target_sequences)

# Create a DataLoader for the dataset
dataloader = torch.utils.data.DataLoader(dataset, batch_size=BATCH_SIZE, shuffle=True)

# Set the device
device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')

model = RephraseGenerator(vocab_size=1000, embedding_dim=256, hidden_size=512, num_layers=2, dropout=0.2)

# Move the model to the device
model.to(device)

# Set the optimizer and loss function
optimizer = optim.AdamW(model.parameters())
loss_fn = nn.CrossEntropyLoss()

train(model, dataloader, optimizer, device)

train函数

# Training loop
def train(model, data_loader, optimizer, device):
    model.train()
    epoch_loss = 0
    for input_sequence, target_sequence in data_loader:
        input_sequence = input_sequence.to(device)
        target_sequence = target_sequence.to(device)
        optimizer.zero_grad()
        predictions = model(input_sequence, target_sequence)
        loss = rephrase_loss(predictions, target_sequence)
        loss.backward()
        optimizer.step()
        epoch_loss += loss.item()
    return epoch_loss / len(data_loader)

ensure_tensor_size函数

def ensure_tensor_size(tensor_list, size):
    """Ensures that each tensor in the list has the given size.
    
    If a tensor has a different size, it is padded with zeros.
    
    Args:
        tensor_list: a list of tensors
        size: an integer representing the desired size of the tensors
    
    Returns:
        a new list of tensors with the same size
    """
    padded_tensor_list = []
    for tensor in tensor_list:
        if tensor.size(0) < size:
            tensor = F.pad(tensor, (0, size - tensor.size(0)), value=0)
        elif tensor.size(0) > size:
            tensor = tensor[:size]
        padded_tensor_list.append(tensor)
    return padded_tensor_list

错误栈信息

IndexError                                Traceback (most recent call last)
<ipython-input-5-1e843ae1e696> in <module>
    233 loss_fn = nn.CrossEntropyLoss()
    234 
--> 235 train(model, dataloader, optimizer, device)

5 frames
/usr/local/lib/python3.8/dist-packages/torch/nn/functional.py in embedding(input, weight, padding_idx, max_norm, norm_type, scale_grad_by_freq, sparse)
   2208         # remove once script supports set_grad_enabled
   2209         _no_grad_embedding_renorm_(weight, input, max_norm, norm_type)
-> 2210     return torch.embedding(weight, input, padding_idx, scale_grad_by_freq, sparse)
   2211 
   2212 

IndexError: index out of range in self

数据样例

[('Outdoor toys for kids',
  [" Kids' outdoor toys",
   ' Outdoor playthings for children',
   " Children's outdoor entertainment",
   ' Outdoor games for young ones ']),
 ('ducational toys for kids',
  [" Kids' educational toys",
   ' Educational playthings for children',
   " Children's educational entertainment",
   ' Educational games for young ones']),
 ('Dolls for girls',
  [" Girls' dolls",
   ' Dolls for little girls',
   ' Entertainment for young girls',
   ' Toys for young girls']),
 ("Kids' swings",
  [" Children's swings", ' Swings for kids', ' Swings for young ones', '']),
 ("Kids' footballs",
  [" Children's footballs",
   ' Footballs for kids',
   ' Footballs for young ones',
   ' Entertainment for young kids']),
 ('Pogo sticks for children',
  [" Kids' pogo sticks",
   " Children's pogo sticks",
   ' Pogo sticks for kids',
   ' Pogo sticks for young ones']),
 ('Holiday gifts',
  [' Gifts for the holidays',
   ' Gifts for special occasions',
   ' Gifts for celebrations',
   ' Gifts for loved ones']),
 ('Best holiday gifts',
  [' Top holiday gifts',
   ' Highly rated holiday gifts',
   ' Recommended holiday gifts',
   ' Best gifts for the holidays']),
 ('Popular holiday gifts',
  [' Best-selling holiday gifts',
   ' Most sought-after holiday gifts',
   ' Trending holiday gifts',
   ' Hot holiday gifts']),
 ('Holiday gifts for kids',
  [' Gifts for children during the holidays',
   ' Gifts for young ones during the holidays',
   ' Gifts for little ones during the holidays',
   ' Gifts for minors during the holidays']),
 ('Holiday gifts for men',
  [' Gifts for men during the holidays',
   ' Gifts for him during the holidays',
   ' Gifts for fathers during the holidays',
   ' Gifts for husbands during the holidays']),
 ('Holiday gifts for women',
  [' Gifts for women during the holidays',
   ' Gifts for her during the holidays',
   ' Gifts for mothers during the holidays',
   ' Gifts for wives during the holidays']),
 ('Holiday gifts for teens',
  [' Gifts for teenagers during the holidays',
   ' Gifts for adolescents during the holidays',
   ' Gifts for young adults during the holidays',
   ' Gifts for older kids during the holidays']),
 ('Holiday gifts for parents',
  [' Gifts for parents during the holidays',
   ' Gifts for mom and dad during the holidays',
   ' Gifts for caregivers during the holidays',
   ' Gifts for adults during the holidays']),
 ('Holiday gifts for grandparents',
  [' Gifts for grandparents during the holidays',
   ' Gifts for grandpa and grandma during the holidays',
   ' Gifts for senior citizens during the holidays',
   ' Gifts for older adults during the holidays']),
 ('Holiday gifts for friends',
  [' Gifts for friends during the holidays',
   ' Gifts for close friends during the holidays',
   ' Gifts for companions during the holidays',
   ' Gifts for peers during the holidays']),
 ('Holiday gifts for coworkers',
  [' Gifts for coworkers during the holidays',
   ' Gifts for colleagues during the holidays',
   ' Gifts for associates during the holidays',
   ' Gifts for professionals during the holidays']),
 ('Holiday gifts for pets',
  [' Gifts for pets during the holidays',
   ' Gifts for dogs during the holidays',
   ' Gifts for cats during the holidays',
   ' Gifts for animals during the holidays']),
 ('Holiday gifts for gamers',
  [' Gifts for gamers during the holidays',
   ' Gifts for video game enthusiasts during the holidays',
   ' Gifts for console gamers during the holidays',
   ' Gifts for PC gamers during the holidays']),
 ('Holiday gifts for hikers',
  [' Gifts for hikers during the holidays',
   ' Gifts for outdoor enthusiasts during the holidays',
   ' Gifts for walkers during the holidays',
   ' Gifts for nature lovers during the holidays']),
 ('Holiday gifts for book lovers',
  [' Gifts for book lovers during the holidays',
   ' Gifts for readers during the holidays',
   ' Gifts for bibliophiles during the holidays',
   ' Gifts for literature enthusiasts during the holidays']),
 ('Holiday gifts for foodies',
  [' Gifts for foodies during the holidays',
   ' Gifts for gourmet cooks during the holidays',
   ' Gifts for culinary enthusiasts during the holidays',
   ' Gifts for epicures during the holidays']),
 ('Holiday gifts for knitters and crocheters',
  [' Gifts for knitters and crocheters during the holidays',
   ' Gifts for fiber artists during the holidays',
   ' Gifts for yarn enthusiasts during the holidays',
   ' Gifts for needlework enthusiasts during the holidays']),
 ('Holiday gifts for sewers and quilters',
  [' Gifts for sewers and quilters during the holidays',
   ' Gifts for needleworkers during the holidays',
   ' Gifts for seamstresses during the holidays',
   ' Gifts for tailors during the holidays']),
 ('Holiday gifts for DIYers',
  [' Gifts for DIYers during the holidays',
   ' Gifts for home improvement enthusiasts during the holidays',
   ' Gifts for handymen and handywomen during the holidays',
   ' Gifts for crafters during the holidays']),
 ('Holiday gifts for mechanics',
  [' Gifts for mechanics during the holidays',
   ' Gifts for auto mechanics during the holidays',
   ' Gifts for mechanic enthusiasts during the holidays',
   ' Gifts for technicians during the holidays']),
 ('Holiday gifts for handymen and handywomen',
  [' Gifts for handymen and handywomen during the holidays',
   ' Gifts for DIY enthusiasts during the holidays',
   ' Gifts for home improvement experts during the holidays',
   ' Gifts for craftspeople during the holidays']),
 ('Luggage sets',
  [' Best luggage',
   ' Travel luggage',
   ' Suitcase sets',
   ' Best luggage sets ']),
 ('Travel backpacks',
  [' Backpacks for travel',
   ' Best travel backpacks',
   ' Popular travel backpacks',
   ' Backpacks for vacation ']),
 ('Travel pillows',
  [' Best travel pillows',
   ' Top-rated travel pillows',
   ' Popular travel pillows',
   ' Best pillows for travel ']),
 ('Travel neck pillows',
  [' Best travel neck pillows',
   ' Best-selling travel neck pillows',
   ' Neck pillows for travel',
   ' Popular travel neck pillows '])]

问题排查与修复方案

核心问题1:词汇表大小不匹配

GPT2Tokenizer的内置词汇表大小为50257,但初始化RephraseGenerator时硬编码了vocab_size=1000,导致模型的embedding层权重矩阵只有1000个词向量。而tokenizer编码后的token索引普遍远大于1000,输入embedding层时直接触发索引越界错误。

修复:用tokenizer的实际词汇表大小初始化模型:

model = RephraseGenerator(vocab_size=tokenizer.vocab_size, embedding_dim=256, hidden_size=512, num_layers=2, dropout=0.2)

核心问题2:目标序列处理逻辑错误

数据样例中每个query对应多个改写句,但直接把整个rephrases列表传给tokenizer.encode(该函数仅接受字符串输入),会导致编码失败,生成无效的token索引,进一步触发越界。

修复:遍历每个改写句,生成一对一的输入-目标对,同时过滤空字符串:

input_sequences = []
target_sequences = []
for query, rephrases in data:
    for rephrase in rephrases:
        cleaned_rephrase = rephrase.strip()
        if cleaned_rephrase:  # 跳过空的改写句
            input_sequences.append(query)
            target_sequences.append(cleaned_rephrase)

核心问题3:手动padding/truncation的不合理性

  1. padding值错误:用0作为padding值,但GPT2的<|endoftext|>索引是50256,0是普通token索引,用0填充会被模型视为有效词汇,同时如果模型vocab_size设置错误,0也可能触发越界。
  2. 长度设置不合理:硬编码截断/ padding到长度4,会丢失绝大多数句子的有效信息,截断后语义完全破坏。

修复:使用tokenizer的批量编码功能,自动处理padding和截断,并指定正确的padding token:

# GPT2默认没有pad token,指定eos token作为pad token
tokenizer.pad_token = tokenizer.eos_token
max_seq_length = 32  # 根据你的数据实际长度调整

# 批量编码输入和目标序列
encoded_inputs = tokenizer(
    input_sequences,
    padding='max_length',
    truncation=True,
    max_length=max_seq_length,
    return_tensors='pt'
)
encoded_targets = tokenizer(
    target_sequences,
    padding='max_length',
    truncation=True,
    max_length=max_seq_length,
    return_tensors='pt'
)

# 获取处理后的tensor
input_sequences = encoded_inputs['input_ids']
target_sequences = encoded_targets['input_ids']

额外检查点

  1. 确认RephraseDataset类正确返回每个样本的input和target tensor,确保DataLoader能正常批量加载。
  2. 检查rephrase_loss函数的实现,确保predictions的形状为[batch_size, seq_len, vocab_size],target_sequence的形状为[batch_size, seq_len],符合CrossEntropyLoss的输入要求。

内容的提问来源于stack exchange,提问作者mchd

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 14:15:41