如何高效判断DataFrame单元格值包含关系并填充目标列?
高效实现DataFrame单字词匹配多字词的解决方案
问题背景
现有包含“单字词”和“多字词”两列的DataFrame,需新增第三列,填充所有包含对应“单字词”的“多字词”内容(最多保留5个),同时处理<none>的特殊情况(返回空字符串)。
示例数据
import pandas as pd # 构造示例DataFrame data = { "单字词": ["Bird", "Stone", "Blood", "<none>"], "多字词": ["Bird with no blood", "Stone that killed the bird", "Bird without brains", "stone and blood"] } df = pd.DataFrame(data)
高效解决方案
摒弃嵌套循环,通过预存多字词列表+列表推导+apply实现,效率远高于手动嵌套循环,且代码简洁易读:
# 预提取所有多字词及其小写版本,避免重复计算 all_phrases = df["多字词"].tolist() all_phrases_lower = [phrase.lower() for phrase in all_phrases] def get_matching_phrases(word): # 处理特殊值<none> if word == "<none>": return "" # 统一转为小写,实现不区分大小写匹配 target_word = word.lower() # 筛选匹配的多字词,最多保留5个 matches = [all_phrases[i] for i, p in enumerate(all_phrases_lower) if target_word in p] # 用逗号拼接结果 return ", ".join(matches[:5]) # 新增目标列 df["含单字词的多字词"] = df["单字词"].apply(get_matching_phrases)
结果验证
运行后得到的DataFrame与预期一致:
| 单字词 | 多字词 | 含单字词的多字词 |
|---|---|---|
| Bird | Bird with no blood | Bird with no blood, Stone that killed the bird, Bird without brains |
| Stone | Stone that killed the bird | Stone that killed the bird, stone and blood |
| Blood | Bird without brains | Bird with no blood, stone and blood |
| stone and blood |
效率说明
- 预存数据:提前计算多字词列表及其小写版本,避免每次匹配时重复转换和读取DataFrame,减少IO开销
- 列表推导:Python内置的列表推导是底层优化的循环,比手动嵌套for循环效率提升数倍
- 按需匹配:仅对每个单字词遍历一次多字词列表,时间复杂度为O(n*m)(n为单字词行数,m为多字词总数),远优于低效的嵌套循环实现
内容的提问来源于stack exchange,提问作者Павел Прахов
相关产品推荐
相关产品推荐

