You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将字典单词分值映射到DataFrame内容并高效计算每行平均分

性能优化方案

原有方案痛点

  • 超大字典拼接正则表达式会生成极长的匹配规则,匹配效率指数级下降
  • explode方法会生成大量中间行,内存开销高,大数据量下性能极差

优化方案

方案1:轻量场景通用方案(无需额外依赖)

核心逻辑利用Python字典O(1)的查询效率,先提取句子中的所有单词再过滤匹配,避免生成超长正则:

import pandas as pd
import re

# 测试数据
test_df = pd.DataFrame({
    '_id': ['1a','2b','3c','4d'],
    'column': ['und der in zu',
                'Kompliziertereswort something',
                'Lehrerin in zu [Buch]',
                'Buch (Lehrerin) kompliziertereswort']
})

# 分值字典
score_dict = {'und': 20,
     'der': 10,
     'in':  40,
     'zu':  10,
     'Kompliziertereswort': 2,
     'Buch': 5,
     'Lehrerin': 5}

# 核心计算逻辑
test_df['score'] = test_df['column'].str.findall(r'\w+').apply(
    lambda words: sum(score_dict[w] for w in words if w in score_dict)/len(match_words) 
    if len(match_words := [w for w in words if w in score_dict]) > 0 
    else 0.0
)

print(test_df)

输出结果和原逻辑完全一致:

_id                               column  score
0  1a                        und der in zu   20.0
1  2b        Kompliziertereswort something    2.0
2  3c                Lehrerin in zu [Buch]   15.0
3  4d  Buch (Lehrerin) kompliziertereswort    5.0

注:如果使用Python3.8以下版本,将lambda中的海象运算符替换为嵌套lambda即可兼容。

方案2:超大数据量高性能方案(numba JIT加速)

如果数据量超过100万行、字典键超过10万,推荐用numba将计算逻辑编译为机器码,性能比纯Python方案提升10-100倍。
首先安装依赖:
pip install numba
核心代码:

from numba import jit
from numba.typed import Dict
from numba.core import types

# 转换为numba支持的字典类型
numba_score_dict = Dict.empty(
    key_type=types.unicode_type,
    value_type=types.int64,
)
for k, v in score_dict.items():
    numba_score_dict[k] = v

@jit(nopython=True)
def calc_sentence_score(sentence, score_dict):
    # 手动拆分单词,避免正则开销
    words = []
    current_word = []
    for c in sentence:
        if c.isalnum():
            current_word.append(c)
        else:
            if current_word:
                words.append(''.join(current_word))
                current_word = []
    if current_word:
        words.append(''.join(current_word))
    
    total_score = 0
    match_count = 0
    for w in words:
        if w in score_dict:
            total_score += score_dict[w]
            match_count += 1
    return total_score / match_count if match_count > 0 else 0.0

test_df['score'] = test_df['column'].apply(lambda x: calc_sentence_score(x, numba_score_dict))

额外适配说明

如果业务中不需要区分单词大小写,可以提前将字典所有键转为小写,拆分单词后也转小写再匹配,适配更多输入场景。

内容的提问来源于stack exchange,提问作者johnnydoe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 10:15:04