You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于自定义列表从Pandas DataFrame中提取指定字符串

Pandas 按自定义列表提取字符串生成新列

原始数据

import pandas as pd
df = pd.DataFrame({'a': ['hair color other family, friends ', 'family, friends hair color']})

对应的DataFrame:

索引a
0hair color other family, friends
1family, friends hair color

自定义提取列表

items = ['hair color', 'other', 'family, friends']

期望输出

import numpy as np
desired_output = pd.DataFrame({
    'a': ['hair color other family, friends ', 'family, friends hair color'],
    'hair color': ['hair color', 'hair color'],
    'other': ['other', np.nan],
    'family, friends': ['family, friends', 'family, friends']
})

对应的DataFrame:

索引ahair colorotherfamily, friends
0hair color other family, friendshair colorotherfamily, friends
1family, friends hair colorhair colorNaNfamily, friends

解决方案

直接通过遍历自定义列表,结合字符串包含检查生成新列:

import pandas as pd
import numpy as np

df = pd.DataFrame({'a': ['hair color other family, friends ', 'family, friends hair color']})
items = ['hair color', 'other', 'family, friends']

for item in items:
    df[item] = df['a'].apply(lambda x: item if item in x else np.nan)

# 查看结果
print(df)

说明

  • 对每个自定义元素,创建同名新列
  • 用apply遍历每行数据,判断元素是否存在于当前字符串中:存在则返回元素本身,不存在则填充NaN

内容的提问来源于stack exchange,提问作者johnjohn

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 02:10:37