解决pandas报错single positional indexer is out-of-bounds 跳过df3缺失词汇计算得分
问题解决
报错原因
当句子中出现的词汇不在df3的Word列范围内时,df3[df3.Word == word]会返回空DataFrame,此时调用.iloc[0]取第一行数据的操作就会触发IndexError: single positional indexer is out-of-bounds报错。
解决方案
方案1:最小改动原有代码
只需在查询属性前先判断匹配结果是否为空,为空直接跳过当前词即可:
pos_score = 0 neg_score = 0 for i in range(len(filtered_sentence)): for word in filtered_sentence[i]: # 先查询匹配结果 match_result = df3[df3.Word == word] # 无匹配结果直接跳过当前词 if len(match_result) == 0: continue # 有匹配结果再走原有逻辑 if match_result.iloc[0]['Negative'] == 2009: neg_score = neg_score + 1 elif match_result.iloc[0]['Positive'] == 2009: pos_score = pos_score + 1 else: break
方案2:性能优化版(推荐)
循环中反复过滤DataFrame效率极低,提前把df3转为字典映射,单次查询时间复杂度直接降到O(1),数据量越大优势越明显:
pos_score = 0 neg_score = 0 # 提前构建词-属性映射字典 word_attr_map = df3.set_index('Word').to_dict('index') for i in range(len(filtered_sentence)): for word in filtered_sentence[i]: # 词不在映射中直接跳过 if word not in word_attr_map: continue attr = word_attr_map[word] if attr['Negative'] == 2009: neg_score += 1 elif attr['Positive'] == 2009: pos_score += 1 else: break
内容的提问来源于stack exchange,提问作者Shreyanshu
相关产品推荐
相关产品推荐

