You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用SBERT生成Embeddings时触发IndexError索引越界问题排查

SBERT encode嵌套列表触发IndexError的排查思路

问题场景

我有如下嵌套列表:

list1=[['brute-force',
  'password-guessing',
  'password-guessing',
  'default-credentials',
  'shell'],
 ['malware',
  'ddos',
  'phishing',
  'spam',
  'botnet',
  'cryptojacking',
  'xss',
  'sqli',
  'vulnerability'],
 ['sensitive-information']]

调用SBERT的encode方法:

embeddings1 = sbert_model.encode(list1, convert_to_tensor=True)

触发如下错误:

IndexError                                Traceback (most recent call last)
~\AppData\Local\Temp/ipykernel_16484/3954167634.py in <module>
----> 1 embeddings2 = sbert_model.encode(list3, convert_to_tensor=True)

~\anaconda3\envs\tensorflow_env\lib\site-packages\sentence_transformers\SentenceTransformer.py in encode(self, sentences, batch_size, show_progress_bar, output_value, convert_to_numpy, convert_to_tensor, device, normalize_embeddings)
    159         for start_index in trange(0, len(sentences), batch_size, desc="Batches", disable=not show_progress_bar):
    160             sentences_batch = sentences_sorted[start_index:start_index+batch_size]
--> 161             features = self.tokenize(sentences_batch)
    162             features = batch_to_device(features, device)
    163 

~\anaconda3\envs\tensorflow_env\lib\site-packages\sentence_transformers\SentenceTransformer.py in tokenize(self, texts)
    317         Tokenizes the texts
    318         """
--> 319         return self._first_module().tokenize(texts)
    320 
    321     def get_sentence_features(self, *features):

~\anaconda3\envs\tensorflow_env\lib\site-packages\sentence_transformers\models\Transformer.py in tokenize(self, texts)
    101             for text_tuple in texts:
    102                 batch1.append(text_tuple[0])
--> 103                 batch2.append(text_tuple[1])
    104             to_tokenize = [batch1, batch2]
    105 

IndexError: list index out of range

排查思路

  • 输入格式不匹配:SBERT的encode默认接受一维文本列表(如["text1", "text2"]),你的二维嵌套列表会被模型识别为句对输入结构(即每个子列表需包含两个配对文本,用于句对相似度计算)。但第三个子列表只有1个元素,模型尝试取第二个元素时触发索引越界。
  • 调整输入结构:若需求是为每个单独文本生成embedding,先将嵌套列表扁平化:
    flat_list = [item for sublist in list1 for item in sublist]
    embeddings1 = sbert_model.encode(flat_list, convert_to_tensor=True)
    
  • 确认句对输入合法性:若确实需要处理句对,需保证每个子列表都包含两个元素,补全缺失内容或调整输入结构。
  • 检查变量名一致性:错误栈中调用的是list3,但你提供的变量是list1,确认代码中变量名是否统一,避免因变量未定义或结构不符引发隐性问题。

内容的提问来源于stack exchange,提问作者xavi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 16:40:26