You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

递归查找并替换.md文件指定区间内的U+202F窄不换行空格

问题需求

递归查找当前目录下所有.md文件,找出在\begin{document}与\end{document}标记之间(可跨行)包含Unicode字符U+202F(窄不换行空格)的文件,并将该字符替换为普通空格。

用户尝试用Python正则提取标记间文本时触发转义错误:re.error: bad escape \e at position 27,相关代码及报错信息如下:

def finds_files_whose_contents_match_a_regex(filename):
    textfile = open(filename, 'r')
    filetext = textfile.read()
    textfile.close()
    matches = re.findall("\\begin{document}\\s*(.*?)\\s*\\end{document}", filetext)

for root, dirs, files in os.walk("."):
    for filename in files:
        if filename.endswith(".md"):
            filename=os.path.join(root, filename)
            finds_files_whose_contents_match_a_regex(filename)

报错回溯:

Traceback (most recent call last):
  File "./test-bis.py", line 14, in <module>
    finds_files_whose_contents_match_a_regex(filename)
  File "./test-bis.py", line 8, in finds_files_whose_contents_match_a_regex
    matches = re.findall("\\begin{document}\\s*(.*?)\\s*\\end{document}", filetext)
  File "/usr/lib64/python3.10/re.py", line 240, in findall
    return _compile(pattern, flags).findall(string)
  File "/usr/lib64/python3.10/re.py", line 303, in _compile
    p = sre_compile.compile(pattern, flags)
  File "/usr/lib64/python3.10/sre_compile.py", line 788, in compile
    p = sre_parse.parse(p, flags)
  File "/usr/lib64/python3.10/sre_parse.py", line 955, in parse
    p = _parse_sub(source, state, flags & SRE_FLAG_VERBOSE, 0)
  File "/usr/lib64/python3.10/sre_parse.py", line 444, in _parse_sub
    itemsappend(_parse(source, state, verbose, nested + 1,
  File "/usr/lib64/python3.10/sre_parse.py", line 526, in _parse
    code = _escape(source, this, state)
  File "/usr/lib64/python3.10/sre_parse.py", line 427, in _escape
    raise source.error("bad escape %s" % escape, len(escape))
re.error: bad escape \e at position 27
错误原因
  1. 转义冲突:双引号包裹的正则字符串中,\\会被Python解析为单个\,导致\e(来自\\end中的\\e)被当成无效转义序列——Python正则语法里没有\e这个合法转义字符。
  2. 其他问题:原代码未导入re和os模块;正则默认不跨行匹配(.无法匹配换行符);仅提取内容未实现替换和写回逻辑。
修正后的完整代码
import os
import re

def replace_narrow_non_breaking_space(filename):
    # 读取文件内容,指定utf-8编码确保Unicode字符正常读写
    with open(filename, 'r', encoding='utf-8') as f:
        filetext = f.read()
    
    # 用原始字符串定义正则,避免转义冲突;添加re.DOTALL实现跨行匹配
    pattern = r'\\begin{document}\s*(.*?)\s*\\end{document}'
    
    # 自定义替换逻辑:仅替换标记内的U+202F为普通空格
    def process_match(match):
        inner_content = match.group(1)
        cleaned_content = inner_content.replace("\u202F", " ")
        return f'\\begin{document} {cleaned_content} \\end{document}'
    
    # 执行替换
    updated_text = re.sub(pattern, process_match, filetext, flags=re.DOTALL)
    
    # 内容有变化才写回,避免无意义修改
    if updated_text != filetext:
        with open(filename, 'w', encoding='utf-8') as f:
            f.write(updated_text)
        print(f"已处理文件: {filename}")

# 递归遍历当前目录下所有.md文件
for root, _, files in os.walk("."):
    for file in files:
        if file.endswith(".md"):
            full_path = os.path.join(root, file)
            replace_narrow_non_breaking_space(full_path)
关键改进点
  • 使用**原始字符串r''**定义正则,彻底避免Python对反斜杠的二次转义
  • 添加re.DOTALL标志,让.可以匹配换行符,实现跨行匹配标记间的内容
  • 用re.sub结合自定义函数,精准替换标记范围内的目标字符
  • 用with语句管理文件,自动处理文件关闭,避免资源泄漏
  • 显式指定encoding='utf-8',确保U+202F等Unicode字符正确读写
  • 检查内容变化后再写回文件,减少不必要的磁盘操作

内容的提问来源于stack exchange,提问作者Denis Bitouzé

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 20:51:35