You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

根据匹配组长度替换文本:Python正则表达式替换问题求助

问题

背景

我是一名正在学习《Automate the Boring Stuff with Python》的纯新手,想实现如下文本替换:

输入字符串

"Agent Alice tells Agent Bob secret info about Agent Charlie."  

期望输出

"Agent A**** tells Agent B** secret info about Agent C******."

注意事项

  • 星号的数量取决于姓名的长度。

已尝试的代码

import re
concealName = re.compile(r'(Agent)(\s)(\w)(\w)*')
sentence = "Agent Alice tells Agent Bob secret info about Agent Charlie."  
concealName.sub(r'\1\2\3' + '*' * len(r'\4'), sentence)

实际输出

'Agent A** gave Agent B** info about Agent C**'

疑问

我的思路是计算每个匹配的第4组的长度,用该长度乘以'*'生成对应数量的星号,但结果不符合预期,完全不清楚原因及解决方法。


解决方法

你的代码问题出在len(r'\4')这里——r'\4'是固定字符串,长度永远是2,所以不管匹配到多长的姓名后缀,都会生成2个星号,这就是所有结果都是两个星号的原因。

要实现根据匹配内容动态生成星号,需要用替换函数而非字符串拼接。re.sub()允许传入一个函数作为替换参数,这个函数会接收每个匹配对象,你可以在函数里处理每个分组的实际内容,计算所需星号数量。

正确代码示例

import re

def conceal_match(match_obj):
    agent_part = match_obj.group(1)
    space = match_obj.group(2)
    first_char = match_obj.group(3)
    rest_chars = match_obj.group(4) or ""  # 兼容姓名仅含首字母的极端情况
    stars = '*' * len(rest_chars)
    return f"{agent_part}{space}{first_char}{stars}"

# 把正则里的(\w)*改成(\w+),确保匹配首字母后的所有字符
concealName = re.compile(r'(Agent)(\s)(\w)(\w+)')
sentence = "Agent Alice tells Agent Bob secret info about Agent Charlie."  
result = concealName.sub(conceal_match, sentence)
print(result)

代码说明

  1. 正则调整:将(\w)*改为(\w+),这样分组4会匹配首字母后的全部字符(而非每次匹配单个字符、最终仅保留最后一个);如果坚持用*,则需要通过or ""处理分组4为空的情况。
  2. 替换函数:通过match_obj.group()获取每个分组的实际内容,计算剩余字符长度后生成对应星号,最后拼接成替换后的字符串。

运行代码后会得到预期输出:

"Agent A**** tells Agent B** secret info about Agent C******."

内容的提问来源于stack exchange,提问作者user1020500

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 14:20:21