如何替换pandas DataFrame列中与另一列值匹配的内容
错误原因
你遇到的报错本质是Series.str.replace()的参数限制:这个方法的匹配模式参数只接受字符串或者编译好的正则对象,你直接把整列to_hide(Series类型)传进去,pandas没法把可变的Series作为哈希键完成匹配,才会抛出TypeError: 'Series' objects are mutable, thus they cannot be hashed。
这个方法本身不支持逐行动态传入不同的匹配规则,所以你原来的写法逻辑上就行不通。
可行实现
写法1:逐行替换(可读性最好)
用apply指定按行遍历,每一行取当前行的to_hide值,替换对应something里的匹配文本即可:
import pandas as pd df = pd.DataFrame({ 'to_hide':['bla','sth','sth else'], 'something':['there is bla over there','i cannot see anything','sth else is beautiful'] }) df = df.assign( something = df.apply(lambda x: x['something'].replace(x['to_hide'], 'to_hide'), axis=1) )
运行后输出完全符合你的预期:
to_hide something 0 bla there is to_hide over there 1 sth i cannot see anything 2 sth else to_hide is beautiful
这里用Python原生字符串的
replace而不是pandas的str.replace,好处是不会触发正则解析,就算to_hide里有*、.这类正则特殊字符,也不会出现匹配错误。
写法2:列表推导式(大数据量优先选)
如果你的数据行数比较多,apply的遍历效率不算高,用列表推导式速度能快2-3倍,逻辑完全一致:
df['something'] = [s.replace(pat, 'to_hide') for pat, s in zip(df['to_hide'], df['something'])]
补充说明
- 如果需要做正则匹配替换,把原生字符串的
replace换成re.sub就行,提前导入re库定义匹配规则 - 默认是全局替换所有匹配内容,如果只需要替换第一次出现的文本,给
replace加count=1参数即可
内容的提问来源于stack exchange,提问作者corianne1234
相关产品推荐
相关产品推荐

