You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

运行GloVe词向量加载代码触发ValueError,如何跳过错误继续处理下一个词?

报错原因

这个报错是因为GloVe 840B数据集中存在少量特殊词条,这些词条本身包含空格/特殊符号,用split()切割整行后,非数字内容会混入本应是向量值的values[1:]部分,导致转float失败。

解决方案

用try-except捕获对应行的转换错误,出错后直接跳过当前词条继续处理下一行即可,修复后的代码如下:

# 补充遗漏的numpy导入
import numpy as np
GLOVE_DATASET_PATH = 'glove.840B.300d.txt'

from tqdm import tqdm
import string
embeddings_index = {}
f = open(GLOVE_DATASET_PATH, encoding="utf8")
word_counter = 0
error_counter = 0 # 可选:统计出错的行数
for line in tqdm(f):
  values = line.split()
  word = values[0]
  if word in dictionary:
    try:
        coefs = np.asarray(values[1:], dtype='float32')
        embeddings_index[word] = coefs
    except ValueError:
        error_counter += 1
        continue # 出错直接跳过当前行
  word_counter += 1
f.close()

print('Found %s word vectors matching enron data set.' % len(embeddings_index))
print('Total words in GloVe data set: %s' % word_counter)
print(f'Skipped {error_counter} invalid lines.') # 可选:打印出错数量

如果需要更精准处理切割错位问题,也可以调整切割逻辑:固定取行末尾300个元素作为向量值,前面的所有元素合并作为词条内容,就可以兼容带空格的特殊词条,不过直接加try-except是实现跳过报错需求的最简单方案。

内容的提问来源于stack exchange,提问作者Louis Wong

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 13:54:05