基于自定义列表从Pandas DataFrame中提取指定字符串
Pandas 按自定义列表提取字符串生成新列
原始数据
import pandas as pd df = pd.DataFrame({'a': ['hair color other family, friends ', 'family, friends hair color']})
对应的DataFrame:
| 索引 | a |
|---|---|
| 0 | hair color other family, friends |
| 1 | family, friends hair color |
自定义提取列表
items = ['hair color', 'other', 'family, friends']
期望输出
import numpy as np desired_output = pd.DataFrame({ 'a': ['hair color other family, friends ', 'family, friends hair color'], 'hair color': ['hair color', 'hair color'], 'other': ['other', np.nan], 'family, friends': ['family, friends', 'family, friends'] })
对应的DataFrame:
| 索引 | a | hair color | other | family, friends |
|---|---|---|---|---|
| 0 | hair color other family, friends | hair color | other | family, friends |
| 1 | family, friends hair color | hair color | NaN | family, friends |
解决方案
直接通过遍历自定义列表,结合字符串包含检查生成新列:
import pandas as pd import numpy as np df = pd.DataFrame({'a': ['hair color other family, friends ', 'family, friends hair color']}) items = ['hair color', 'other', 'family, friends'] for item in items: df[item] = df['a'].apply(lambda x: item if item in x else np.nan) # 查看结果 print(df)
说明
- 对每个自定义元素,创建同名新列
- 用
apply遍历每行数据,判断元素是否存在于当前字符串中:存在则返回元素本身,不存在则填充NaN
内容的提问来源于stack exchange,提问作者johnjohn
相关产品推荐
相关产品推荐

