如何在Pandas DataFrame中用其他列值替换指定占位符?
解决Pandas按行替换字符串占位符的问题
问题场景
现有包含ID、Name、Comment列的Pandas DataFrame,Comment列存在占位符NameTag和IDTag,需要将这两个占位符分别替换为对应行的Name列与ID列的值。示例输入与输出如下:
输入表格
| ID | Name | Comment |
|---|---|---|
| A1 | Alex | the name NameTag belonging to IDTag is a common name |
| A2 | Alice | Judging the NameTag of IDTag, I feel its a girl's name |
期望输出表格
| ID | Name | Comment |
|---|---|---|
| A1 | Alex | the name Alex belonging to A1 is a common name |
| A2 | Alice | Judging the Alice of A2, I feel its a girl's name |
尝试使用pandas replace函数实现时触发报错:ValueError: Series.replace cannot use dict-value and non-None to_replace。
报错原因
Series.replace()的字典参数仅支持全局固定值替换(比如{'old':'new'}将所有匹配old的字符串替换为new),但这里需要的是按行动态替换(每一行的替换值对应自身的Name/ID,不是全局固定值),直接传递字典的用法不符合该函数的参数规则,因此报错。
解决方案
方法1:用apply逐行处理
这是最直观的方式,逐行对Comment进行两次替换:
import pandas as pd # 构造示例数据 df = pd.DataFrame({ 'ID': ['A1', 'A2'], 'Name': ['Alex', 'Alice'], 'Comment': [ 'the name NameTag belonging to IDTag is a common name', "Judging the NameTag of IDTag, I feel its a girl's name" ] }) # 逐行替换占位符 df['Comment'] = df.apply( lambda row: row['Comment'].replace('NameTag', row['Name']).replace('IDTag', row['ID']), axis=1 )
方法2:用str.replace结合正则与lambda
利用正则匹配占位符,通过lambda函数根据匹配结果动态获取对应行的替换值,适合数据量较大的场景:
df['Comment'] = df['Comment'].str.replace( r'(NameTag|IDTag)', lambda match: df.loc[match.pos, 'Name'] if match.group() == 'NameTag' else df.loc[match.pos, 'ID'], regex=True )
方法3:用字符串格式化
将Comment转换为格式化模板,再逐行填充对应字段:
df['Comment'] = df.apply( lambda row: row['Comment'].replace('NameTag', '{Name}').replace('IDTag', '{ID}').format(**row), axis=1 )
内容的提问来源于stack exchange,提问作者AnalysisChamp
相关产品推荐
相关产品推荐

