字典存在对应Key仍触发KeyError?NLP词向量问题求助
基于词向量的情感分析KeyError问题解决
问题原因
你在构建词汇表字典时搞反了键值对的映射关系:
原本需要以单词为键、对应的索引为值,但你写的代码把索引设为键、单词设为值。这样当你用单词'does'去字典中查找时,自然找不到对应键,触发KeyError。
看这段出问题的代码:
vocab = {a:b for a,b in enumerate(sorted(vocab))}
enumerate(sorted(vocab))生成的是(索引, 单词)的元组,你把索引作为键、单词作为值,导致vocab结构为{0: 'xxx', 1: 'does', ...},而非你需要的{'xxx':0, 'does':1, ...}。
修复方案
将vocab的键值对调换即可:
vocab = {b:a for a,b in enumerate(sorted(vocab))}
调整后vocab以单词为键、索引为值,后续用vocab[t]就能正确获取单词对应的索引,不会再触发KeyError。
原始问题代码
数据加载与分词
txt_fname = 'C:\Users\arune\Desktop\sentiment labelled sentences\amazon_cells_labelled.txt' df = pd.read_table(txt_fname,names=['sentence','sentiment']) df['tokenized'] = df['sentence'].apply(lambda a: word_tokenize(a)) # 构建词汇集合 vocab = set() for tokens in df['tokenized']: for a in tokens: vocab.add(a) len(vocab) # 出问题的vocab字典构建 vocab = {a:b for a,b in enumerate(sorted(vocab))} vocab rand_wv = np.random.rand(len(vocab),300) rand_wv.shape from sklearn.model_selection import train_test_split train, test = train_test_split(df, test_size=0.2,random_state=42)
训练集向量转换代码
X_train = [] for tok_sent in train['tokenized']: doc_vec = np.zeros(300) for t in tok_sent: word_index = vocab[t] word_vec = rand_wv[word_index] doc_vec += word_vec doc_vec = doc_vec/len(tok_sent) X_train.append(doc_vec) X_train = np.array(X_train) y_train = train['sentiment'] X_train.shape
报错信息
KeyError Traceback (most recent call last) <ipython-input-64-add70741d289> in <module> 3 doc_vec = np.zeros(300) 4 for t in tok_sent: ----> 5 word_index = vocab[t] 6 word_vec = rand_wv[word_index] 7 doc_vec += word_vec KeyError: 'does'
内容的提问来源于stack exchange,提问作者Arun elago
相关产品推荐
相关产品推荐

