如何匹配DataFrame行内单词与字典键计算对应平均分值
完整实现代码
import pandas as pd import re # 原始测试数据 test_df = pd.DataFrame({ '_id': ['1a','2b','3c','4d'], 'column': ['und der in zu', 'Kompliziertereswort something', 'Lehrerin in zu [Buch]', 'Buch (Lehrerin) kompliziertereswort'] }) # 权重映射字典 word_weight = {'und': 20, 'der': 10, 'in': 40, 'zu': 10, 'Kompliziertereswort': 2, 'Buch': 5, 'Lehrerin': 5} # 单行分值计算逻辑 def calc_row_score(text): # 正则提取所有纯字母组成的单词,自动过滤括号、方括号等特殊符号 words = re.findall(r'[a-zA-Z]+', text) # 收集匹配到的权重值 matched_scores = [word_weight[w] for w in words if w in word_weight] # 计算平均值,无匹配项时返回0(可根据需求调整返回值) return sum(matched_scores)/len(matched_scores) if matched_scores else 0 # 批量生成score列 test_df['score'] = test_df['column'].apply(calc_row_score) # 输出结果验证 print(test_df)
运行后输出结果和你给出的预期完全一致:
_id column score 0 1a und der in zu 20.0 1 2b Kompliziertereswort something 2.0 2 3c Lehrerin in zu [Buch] 15.0 3 4d Buch (Lehrerin) kompliziertereswort 5.0
之前实现失败的原因
- 正则匹配方案的问题:没有先提取纯字母单词,直接匹配整行内容会被特殊符号干扰,用
[a-zA-Z]+可以直接提取所有字母序列,自动跳过非字母字符,不需要单独处理括号。 - 遍历行方案的问题:你直接遍历了
values[1]这个字符串对象,Python遍历字符串默认输出单个字符,需要先把字符串提取为单词列表再遍历。
内容的提问来源于stack exchange,提问作者johnnydoe
相关产品推荐
相关产品推荐

