You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何匹配DataFrame行内单词与字典键计算对应平均分值

完整实现代码

import pandas as pd
import re

# 原始测试数据
test_df = pd.DataFrame({
    '_id': ['1a','2b','3c','4d'],
    'column': ['und der in zu',
                'Kompliziertereswort something',
                'Lehrerin in zu [Buch]',
                'Buch (Lehrerin) kompliziertereswort']
})

# 权重映射字典
word_weight = {'und': 20,
 'der': 10,
 'in':  40,
 'zu':  10,
 'Kompliziertereswort': 2,
 'Buch': 5,
 'Lehrerin': 5}

# 单行分值计算逻辑
def calc_row_score(text):
    # 正则提取所有纯字母组成的单词,自动过滤括号、方括号等特殊符号
    words = re.findall(r'[a-zA-Z]+', text)
    # 收集匹配到的权重值
    matched_scores = [word_weight[w] for w in words if w in word_weight]
    # 计算平均值,无匹配项时返回0(可根据需求调整返回值)
    return sum(matched_scores)/len(matched_scores) if matched_scores else 0

# 批量生成score列
test_df['score'] = test_df['column'].apply(calc_row_score)

# 输出结果验证
print(test_df)

运行后输出结果和你给出的预期完全一致:

_id                                column  score
0  1a                         und der in zu   20.0
1  2b         Kompliziertereswort something    2.0
2  3c                 Lehrerin in zu [Buch]   15.0
3  4d  Buch (Lehrerin) kompliziertereswort    5.0

之前实现失败的原因

  • 正则匹配方案的问题:没有先提取纯字母单词,直接匹配整行内容会被特殊符号干扰,用[a-zA-Z]+可以直接提取所有字母序列,自动跳过非字母字符,不需要单独处理括号。
  • 遍历行方案的问题:你直接遍历了values[1]这个字符串对象,Python遍历字符串默认输出单个字符,需要先把字符串提取为单词列表再遍历。

内容的提问来源于stack exchange,提问作者johnnydoe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 18:06:03