You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效判断DataFrame单元格值包含关系并填充目标列?

高效实现DataFrame单字词匹配多字词的解决方案

问题背景

现有包含“单字词”和“多字词”两列的DataFrame,需新增第三列,填充所有包含对应“单字词”的“多字词”内容(最多保留5个),同时处理<none>的特殊情况(返回空字符串)。

示例数据

import pandas as pd

# 构造示例DataFrame
data = {
    "单字词": ["Bird", "Stone", "Blood", "<none>"],
    "多字词": ["Bird with no blood", "Stone that killed the bird", "Bird without brains", "stone and blood"]
}
df = pd.DataFrame(data)

高效解决方案

摒弃嵌套循环,通过预存多字词列表+列表推导+apply实现,效率远高于手动嵌套循环,且代码简洁易读:

# 预提取所有多字词及其小写版本,避免重复计算
all_phrases = df["多字词"].tolist()
all_phrases_lower = [phrase.lower() for phrase in all_phrases]

def get_matching_phrases(word):
    # 处理特殊值<none>
    if word == "<none>":
        return ""
    # 统一转为小写,实现不区分大小写匹配
    target_word = word.lower()
    # 筛选匹配的多字词,最多保留5个
    matches = [all_phrases[i] for i, p in enumerate(all_phrases_lower) if target_word in p]
    # 用逗号拼接结果
    return ", ".join(matches[:5])

# 新增目标列
df["含单字词的多字词"] = df["单字词"].apply(get_matching_phrases)

结果验证

运行后得到的DataFrame与预期一致:

单字词多字词含单字词的多字词
BirdBird with no bloodBird with no blood, Stone that killed the bird, Bird without brains
StoneStone that killed the birdStone that killed the bird, stone and blood
BloodBird without brainsBird with no blood, stone and blood
stone and blood

效率说明

  1. 预存数据:提前计算多字词列表及其小写版本,避免每次匹配时重复转换和读取DataFrame,减少IO开销
  2. 列表推导:Python内置的列表推导是底层优化的循环,比手动嵌套for循环效率提升数倍
  3. 按需匹配:仅对每个单字词遍历一次多字词列表,时间复杂度为O(n*m)(n为单字词行数,m为多字词总数),远优于低效的嵌套循环实现

内容的提问来源于stack exchange,提问作者Павел Прахов

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 09:10:44