如何在Pandas中根据条件提取列字符串生成新列?
问题:基于Pandas DataFrame的条件列生成
给定如下Pandas DataFrame:
import pandas as pd df_test = pd.DataFrame(data=None, columns=['file','comment']) df_test.file = ['file_1', 'file_1_v2', 'file_2', 'file_2_v2', 'file_3', 'file_3_v2'] df_test.comment = ['none: 5', 'Replacing: file_1', 'none', 'Replacing: file_2', 'none', 'Replacing: file_3']
需求:创建一个新列,规则为:
- 若
comment列的字符串以Replacing:开头,提取该字符串后半部分填入新列 - 若不符合上述条件,使用对应行的
file列值填充
期望新列结果:
['file_1', 'file_1', 'file_2', 'file_2', 'file_3', 'file_3']
此前尝试的代码无法满足需求,因为它会对所有含冒号的字符串进行分割:
df_test['comment'].str.extract(r'\s(.*)$', expand=False).fillna(df_test['file'])
解决方案
方法1:str.startswith + np.where
通过条件判断分别处理两类行:
import numpy as np df_test['new_col'] = np.where( df_test['comment'].str.startswith('Replacing: '), df_test['comment'].str.replace('Replacing: ', ''), df_test['file'] )
方法2:锚定开头的正则提取
用^锚定字符串开头,仅匹配以Replacing: 起始的内容,提取后半部分后用file列填充空值:
df_test['new_col'] = df_test['comment'].str.extract(r'^Replacing: (.*)$', expand=False).fillna(df_test['file'])
方法3:条件掩码 + 分割赋值
先初始化新列为file列的值,再对符合条件的行替换为分割后的内容:
mask = df_test['comment'].str.startswith('Replacing: ') df_test['new_col'] = df_test['file'] df_test.loc[mask, 'new_col'] = df_test.loc[mask, 'comment'].str.split('Replacing: ').str[1]
验证结果:
print(df_test['new_col'].tolist()) # 输出:['file_1', 'file_1', 'file_2', 'file_2', 'file_3', 'file_3']
内容的提问来源于stack exchange,提问作者Marcus K.
相关产品推荐
相关产品推荐

