如何将字典单词分值映射到DataFrame内容并高效计算每行平均分
性能优化方案
原有方案痛点
- 超大字典拼接正则表达式会生成极长的匹配规则,匹配效率指数级下降
- explode方法会生成大量中间行,内存开销高,大数据量下性能极差
优化方案
方案1:轻量场景通用方案(无需额外依赖)
核心逻辑利用Python字典O(1)的查询效率,先提取句子中的所有单词再过滤匹配,避免生成超长正则:
import pandas as pd import re # 测试数据 test_df = pd.DataFrame({ '_id': ['1a','2b','3c','4d'], 'column': ['und der in zu', 'Kompliziertereswort something', 'Lehrerin in zu [Buch]', 'Buch (Lehrerin) kompliziertereswort'] }) # 分值字典 score_dict = {'und': 20, 'der': 10, 'in': 40, 'zu': 10, 'Kompliziertereswort': 2, 'Buch': 5, 'Lehrerin': 5} # 核心计算逻辑 test_df['score'] = test_df['column'].str.findall(r'\w+').apply( lambda words: sum(score_dict[w] for w in words if w in score_dict)/len(match_words) if len(match_words := [w for w in words if w in score_dict]) > 0 else 0.0 ) print(test_df)
输出结果和原逻辑完全一致:
_id column score 0 1a und der in zu 20.0 1 2b Kompliziertereswort something 2.0 2 3c Lehrerin in zu [Buch] 15.0 3 4d Buch (Lehrerin) kompliziertereswort 5.0
注:如果使用Python3.8以下版本,将lambda中的海象运算符替换为嵌套lambda即可兼容。
方案2:超大数据量高性能方案(numba JIT加速)
如果数据量超过100万行、字典键超过10万,推荐用numba将计算逻辑编译为机器码,性能比纯Python方案提升10-100倍。
首先安装依赖:pip install numba
核心代码:
from numba import jit from numba.typed import Dict from numba.core import types # 转换为numba支持的字典类型 numba_score_dict = Dict.empty( key_type=types.unicode_type, value_type=types.int64, ) for k, v in score_dict.items(): numba_score_dict[k] = v @jit(nopython=True) def calc_sentence_score(sentence, score_dict): # 手动拆分单词,避免正则开销 words = [] current_word = [] for c in sentence: if c.isalnum(): current_word.append(c) else: if current_word: words.append(''.join(current_word)) current_word = [] if current_word: words.append(''.join(current_word)) total_score = 0 match_count = 0 for w in words: if w in score_dict: total_score += score_dict[w] match_count += 1 return total_score / match_count if match_count > 0 else 0.0 test_df['score'] = test_df['column'].apply(lambda x: calc_sentence_score(x, numba_score_dict))
额外适配说明
如果业务中不需要区分单词大小写,可以提前将字典所有键转为小写,拆分单词后也转小写再匹配,适配更多输入场景。
内容的提问来源于stack exchange,提问作者johnnydoe
相关产品推荐
相关产品推荐

