Python:如何用正则批量替换DataFrame多列指定字符并封装函数?
问题原因与解决方案
问题根源
你当前代码的问题在于重复的字典键会被覆盖:Python字典不允许同一个键重复定义,你写的{'A': r'<','A':r'>','A':r'&'}会被自动处理成仅保留最后一组键值对{'A': r'&'},因此只有&被替换,<和>的规则完全失效。
正确实现方式
1. 定义统一的替换规则字典
先把所有替换规则整理成一个字典,注意正则特殊字符(<、>)需要转义(加\),避免被当成正则语法解析:
replace_rules = { r'&': 'and', r'\<': 'less than', r'\>': 'greater than', r"'": "this is an apostrophe", r'"': 'this is a double quotation' }
2. 单列替换示例
直接对目标列应用完整的替换规则:
import pandas as pd df = pd.DataFrame({'A': ['bat<', 'foo>', 'bait&'], 'B': ['abc', 'bar', 'xyz']}) # 对A列应用所有替换规则 df['A'] = df['A'].replace(replace_rules, regex=True)
执行后df['A']的结果会符合你的预期:['batless than', 'foogreater than', 'baitand']
3. 封装复用函数
把替换逻辑封装成函数,支持同时处理多列:
def apply_special_replacements(df, target_columns): replace_rules = { r'&': 'and', r'\<': 'less than', r'\>': 'greater than', r"'": "this is an apostrophe", r'"': 'this is a double quotation' } for col in target_columns: df[col] = df[col].replace(replace_rules, regex=True) return df # 调用示例:处理A列和新增的Name列等 df = apply_special_replacements(df, ['A', 'Name', 'Col1', 'Col2'])
验证结果
执行上述代码后,你的DataFrame会输出预期结果:
A B 0 batless than abc 1 foogreater than bar 2 baitand xyz
内容的提问来源于stack exchange,提问作者Wajih
相关产品推荐
相关产品推荐

