You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提取DataFrame任意列含指定字符串的内容并新增列存储

跨列提取DataFrame中含指定字符串内容的实现方法

核心实现代码

import pandas as pd
import numpy as np

# 构造样例数据,可替换为你的实际DataFrame
data = {
    "yesterday": ["apple_1", "pear_2", "kiwi_3", "apple_1"],
    "today": ["banana_3", "apple_4", "orange_6", "melon_3"],
    "tomorrow": ["cherry_4", "blueberry_1", "banana_2", "banana_4"]
}
df = pd.DataFrame(data, index=[1,2,3,4])

# 定义目标关键词
target_key = "apple"

# 逐行扫描所有列提取匹配内容
df[target_key] = df.apply(
    lambda row: next((val for val in row if isinstance(val, str) and target_key in val), np.nan),
    axis=1
)

print(df)

逻辑说明

  • 适配约束要求:不需要提前指定搜索列,也不需要预知目标字符串前后的拼接内容,仅做包含匹配
  • 默认规则:一行有多个匹配值时取第一个命中的内容,无匹配时自动填充NaN
  • 可选调整:如果需要保留一行内所有匹配结果,可修改返回逻辑,例如下方代码会将所有匹配值用逗号拼接:
df[target_key] = df.apply(
    lambda row: ",".join([val for val in row if isinstance(val, str) and target_key in val]) or np.nan,
    axis=1
)

注:你给出的期望输出中第4行的apple_2属于笔误,运行上述代码第4行将返回和原数据一致的apple_1

内容的提问来源于stack exchange,提问作者IRK

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 03:54:02