You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python:如何用正则批量替换DataFrame多列指定字符并封装函数?

问题原因与解决方案

问题根源

你当前代码的问题在于重复的字典键会被覆盖:Python字典不允许同一个键重复定义,你写的{'A': r'<','A':r'>','A':r'&'}会被自动处理成仅保留最后一组键值对{'A': r'&'},因此只有&被替换,<和>的规则完全失效。

正确实现方式

1. 定义统一的替换规则字典

先把所有替换规则整理成一个字典,注意正则特殊字符(<、>)需要转义(加\),避免被当成正则语法解析:

replace_rules = {
    r'&': 'and',
    r'\<': 'less than',
    r'\>': 'greater than',
    r"'": "this is an apostrophe",
    r'"': 'this is a double quotation'
}

2. 单列替换示例

直接对目标列应用完整的替换规则:

import pandas as pd

df = pd.DataFrame({'A': ['bat<', 'foo>', 'bait&'],
                   'B': ['abc', 'bar', 'xyz']})

# 对A列应用所有替换规则
df['A'] = df['A'].replace(replace_rules, regex=True)

执行后df['A']的结果会符合你的预期:['batless than', 'foogreater than', 'baitand']

3. 封装复用函数

把替换逻辑封装成函数,支持同时处理多列:

def apply_special_replacements(df, target_columns):
    replace_rules = {
        r'&': 'and',
        r'\<': 'less than',
        r'\>': 'greater than',
        r"'": "this is an apostrophe",
        r'"': 'this is a double quotation'
    }
    for col in target_columns:
        df[col] = df[col].replace(replace_rules, regex=True)
    return df

# 调用示例:处理A列和新增的Name列等
df = apply_special_replacements(df, ['A', 'Name', 'Col1', 'Col2'])

验证结果

执行上述代码后,你的DataFrame会输出预期结果:

A    B
0  batless than  abc
1  foogreater than  bar
2       baitand  xyz

内容的提问来源于stack exchange,提问作者Wajih

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 20:15:35