训练RephraseGenerator模型触发IndexError: index out of range in self错误排查
训练自定义RephraseGenerator模型时触发IndexError: index out of range in self错误
问题描述
训练自定义RephraseGenerator模型时,训练函数抛出「IndexError: index out of range in self」错误,以下是相关代码、错误栈和数据样例,求排查问题。
主代码
# Extract input and target sequences from data list input_sequences = [] target_sequences = [] BATCH_SIZE = 64 data = read_csv('gpt-j-data.csv') for query, rephrases in data: input_sequences.append(query) target_sequences.append(rephrases) # Load the tokenizer tokenizer = GPT2Tokenizer.from_pretrained('gpt2') # Tokenize the input and target sequences input_sequences = [tokenizer.encode(sequence, add_special_tokens=True) for sequence in input_sequences] target_sequences = [tokenizer.encode(sequence, add_special_tokens=True) for sequence in target_sequences] # Convert the input and target sequences to tensors input_sequences = [torch.tensor(sequence) for sequence in input_sequences] target_sequences = [torch.tensor(sequence) for sequence in target_sequences] input_sequences = ensure_tensor_size(input_sequences, 4) target_sequences = ensure_tensor_size(target_sequences, 4) # Create a RephraseDataset object from the input and target sequences dataset = RephraseDataset(input_sequences, target_sequences) # Create a DataLoader for the dataset dataloader = torch.utils.data.DataLoader(dataset, batch_size=BATCH_SIZE, shuffle=True) # Set the device device = torch.device('cuda' if torch.cuda.is_available() else 'cpu') model = RephraseGenerator(vocab_size=1000, embedding_dim=256, hidden_size=512, num_layers=2, dropout=0.2) # Move the model to the device model.to(device) # Set the optimizer and loss function optimizer = optim.AdamW(model.parameters()) loss_fn = nn.CrossEntropyLoss() train(model, dataloader, optimizer, device)
train函数
# Training loop def train(model, data_loader, optimizer, device): model.train() epoch_loss = 0 for input_sequence, target_sequence in data_loader: input_sequence = input_sequence.to(device) target_sequence = target_sequence.to(device) optimizer.zero_grad() predictions = model(input_sequence, target_sequence) loss = rephrase_loss(predictions, target_sequence) loss.backward() optimizer.step() epoch_loss += loss.item() return epoch_loss / len(data_loader)
ensure_tensor_size函数
def ensure_tensor_size(tensor_list, size): """Ensures that each tensor in the list has the given size. If a tensor has a different size, it is padded with zeros. Args: tensor_list: a list of tensors size: an integer representing the desired size of the tensors Returns: a new list of tensors with the same size """ padded_tensor_list = [] for tensor in tensor_list: if tensor.size(0) < size: tensor = F.pad(tensor, (0, size - tensor.size(0)), value=0) elif tensor.size(0) > size: tensor = tensor[:size] padded_tensor_list.append(tensor) return padded_tensor_list
错误栈信息
IndexError Traceback (most recent call last) <ipython-input-5-1e843ae1e696> in <module> 233 loss_fn = nn.CrossEntropyLoss() 234 --> 235 train(model, dataloader, optimizer, device) 5 frames /usr/local/lib/python3.8/dist-packages/torch/nn/functional.py in embedding(input, weight, padding_idx, max_norm, norm_type, scale_grad_by_freq, sparse) 2208 # remove once script supports set_grad_enabled 2209 _no_grad_embedding_renorm_(weight, input, max_norm, norm_type) -> 2210 return torch.embedding(weight, input, padding_idx, scale_grad_by_freq, sparse) 2211 2212 IndexError: index out of range in self
数据样例
[('Outdoor toys for kids', [" Kids' outdoor toys", ' Outdoor playthings for children', " Children's outdoor entertainment", ' Outdoor games for young ones ']), ('ducational toys for kids', [" Kids' educational toys", ' Educational playthings for children', " Children's educational entertainment", ' Educational games for young ones']), ('Dolls for girls', [" Girls' dolls", ' Dolls for little girls', ' Entertainment for young girls', ' Toys for young girls']), ("Kids' swings", [" Children's swings", ' Swings for kids', ' Swings for young ones', '']), ("Kids' footballs", [" Children's footballs", ' Footballs for kids', ' Footballs for young ones', ' Entertainment for young kids']), ('Pogo sticks for children', [" Kids' pogo sticks", " Children's pogo sticks", ' Pogo sticks for kids', ' Pogo sticks for young ones']), ('Holiday gifts', [' Gifts for the holidays', ' Gifts for special occasions', ' Gifts for celebrations', ' Gifts for loved ones']), ('Best holiday gifts', [' Top holiday gifts', ' Highly rated holiday gifts', ' Recommended holiday gifts', ' Best gifts for the holidays']), ('Popular holiday gifts', [' Best-selling holiday gifts', ' Most sought-after holiday gifts', ' Trending holiday gifts', ' Hot holiday gifts']), ('Holiday gifts for kids', [' Gifts for children during the holidays', ' Gifts for young ones during the holidays', ' Gifts for little ones during the holidays', ' Gifts for minors during the holidays']), ('Holiday gifts for men', [' Gifts for men during the holidays', ' Gifts for him during the holidays', ' Gifts for fathers during the holidays', ' Gifts for husbands during the holidays']), ('Holiday gifts for women', [' Gifts for women during the holidays', ' Gifts for her during the holidays', ' Gifts for mothers during the holidays', ' Gifts for wives during the holidays']), ('Holiday gifts for teens', [' Gifts for teenagers during the holidays', ' Gifts for adolescents during the holidays', ' Gifts for young adults during the holidays', ' Gifts for older kids during the holidays']), ('Holiday gifts for parents', [' Gifts for parents during the holidays', ' Gifts for mom and dad during the holidays', ' Gifts for caregivers during the holidays', ' Gifts for adults during the holidays']), ('Holiday gifts for grandparents', [' Gifts for grandparents during the holidays', ' Gifts for grandpa and grandma during the holidays', ' Gifts for senior citizens during the holidays', ' Gifts for older adults during the holidays']), ('Holiday gifts for friends', [' Gifts for friends during the holidays', ' Gifts for close friends during the holidays', ' Gifts for companions during the holidays', ' Gifts for peers during the holidays']), ('Holiday gifts for coworkers', [' Gifts for coworkers during the holidays', ' Gifts for colleagues during the holidays', ' Gifts for associates during the holidays', ' Gifts for professionals during the holidays']), ('Holiday gifts for pets', [' Gifts for pets during the holidays', ' Gifts for dogs during the holidays', ' Gifts for cats during the holidays', ' Gifts for animals during the holidays']), ('Holiday gifts for gamers', [' Gifts for gamers during the holidays', ' Gifts for video game enthusiasts during the holidays', ' Gifts for console gamers during the holidays', ' Gifts for PC gamers during the holidays']), ('Holiday gifts for hikers', [' Gifts for hikers during the holidays', ' Gifts for outdoor enthusiasts during the holidays', ' Gifts for walkers during the holidays', ' Gifts for nature lovers during the holidays']), ('Holiday gifts for book lovers', [' Gifts for book lovers during the holidays', ' Gifts for readers during the holidays', ' Gifts for bibliophiles during the holidays', ' Gifts for literature enthusiasts during the holidays']), ('Holiday gifts for foodies', [' Gifts for foodies during the holidays', ' Gifts for gourmet cooks during the holidays', ' Gifts for culinary enthusiasts during the holidays', ' Gifts for epicures during the holidays']), ('Holiday gifts for knitters and crocheters', [' Gifts for knitters and crocheters during the holidays', ' Gifts for fiber artists during the holidays', ' Gifts for yarn enthusiasts during the holidays', ' Gifts for needlework enthusiasts during the holidays']), ('Holiday gifts for sewers and quilters', [' Gifts for sewers and quilters during the holidays', ' Gifts for needleworkers during the holidays', ' Gifts for seamstresses during the holidays', ' Gifts for tailors during the holidays']), ('Holiday gifts for DIYers', [' Gifts for DIYers during the holidays', ' Gifts for home improvement enthusiasts during the holidays', ' Gifts for handymen and handywomen during the holidays', ' Gifts for crafters during the holidays']), ('Holiday gifts for mechanics', [' Gifts for mechanics during the holidays', ' Gifts for auto mechanics during the holidays', ' Gifts for mechanic enthusiasts during the holidays', ' Gifts for technicians during the holidays']), ('Holiday gifts for handymen and handywomen', [' Gifts for handymen and handywomen during the holidays', ' Gifts for DIY enthusiasts during the holidays', ' Gifts for home improvement experts during the holidays', ' Gifts for craftspeople during the holidays']), ('Luggage sets', [' Best luggage', ' Travel luggage', ' Suitcase sets', ' Best luggage sets ']), ('Travel backpacks', [' Backpacks for travel', ' Best travel backpacks', ' Popular travel backpacks', ' Backpacks for vacation ']), ('Travel pillows', [' Best travel pillows', ' Top-rated travel pillows', ' Popular travel pillows', ' Best pillows for travel ']), ('Travel neck pillows', [' Best travel neck pillows', ' Best-selling travel neck pillows', ' Neck pillows for travel', ' Popular travel neck pillows '])]
问题排查与修复方案
核心问题1:词汇表大小不匹配
GPT2Tokenizer的内置词汇表大小为50257,但初始化RephraseGenerator时硬编码了vocab_size=1000,导致模型的embedding层权重矩阵只有1000个词向量。而tokenizer编码后的token索引普遍远大于1000,输入embedding层时直接触发索引越界错误。
修复:用tokenizer的实际词汇表大小初始化模型:
model = RephraseGenerator(vocab_size=tokenizer.vocab_size, embedding_dim=256, hidden_size=512, num_layers=2, dropout=0.2)
核心问题2:目标序列处理逻辑错误
数据样例中每个query对应多个改写句,但直接把整个rephrases列表传给tokenizer.encode(该函数仅接受字符串输入),会导致编码失败,生成无效的token索引,进一步触发越界。
修复:遍历每个改写句,生成一对一的输入-目标对,同时过滤空字符串:
input_sequences = [] target_sequences = [] for query, rephrases in data: for rephrase in rephrases: cleaned_rephrase = rephrase.strip() if cleaned_rephrase: # 跳过空的改写句 input_sequences.append(query) target_sequences.append(cleaned_rephrase)
核心问题3:手动padding/truncation的不合理性
- padding值错误:用0作为padding值,但GPT2的
<|endoftext|>索引是50256,0是普通token索引,用0填充会被模型视为有效词汇,同时如果模型vocab_size设置错误,0也可能触发越界。 - 长度设置不合理:硬编码截断/ padding到长度4,会丢失绝大多数句子的有效信息,截断后语义完全破坏。
修复:使用tokenizer的批量编码功能,自动处理padding和截断,并指定正确的padding token:
# GPT2默认没有pad token,指定eos token作为pad token tokenizer.pad_token = tokenizer.eos_token max_seq_length = 32 # 根据你的数据实际长度调整 # 批量编码输入和目标序列 encoded_inputs = tokenizer( input_sequences, padding='max_length', truncation=True, max_length=max_seq_length, return_tensors='pt' ) encoded_targets = tokenizer( target_sequences, padding='max_length', truncation=True, max_length=max_seq_length, return_tensors='pt' ) # 获取处理后的tensor input_sequences = encoded_inputs['input_ids'] target_sequences = encoded_targets['input_ids']
额外检查点
- 确认
RephraseDataset类正确返回每个样本的input和target tensor,确保DataLoader能正常批量加载。 - 检查
rephrase_loss函数的实现,确保predictions的形状为[batch_size, seq_len, vocab_size],target_sequence的形状为[batch_size, seq_len],符合CrossEntropyLoss的输入要求。
内容的提问来源于stack exchange,提问作者mchd
相关产品推荐
相关产品推荐

