按指定规则替换普通空格为非断空格的Python实现求助
问题描述
原始需求
需要批量处理目录中的文件,将普通空格按规则替换为非断空格(\u00A0)。例如句子"You need to walk 5 km.",需将数字5与单位km之间的空格替换为非断空格。
现有代码会替换所有符合单位+空格的内容,无法精准定位目标空格:
import os unites = ['km', 'm', 'cm', 'mm', 'mi', 'yd', 'ft', 'in'] # iterate and read all files in the directory for file in os.listdir(): # check if the file is a file if os.path.isfile(file): # open the file with open(file, 'r', encoding='utf-8') as f: # read the file content = f.read() # search for exemple in the file for i in unites: if i in content: # find the next character after the unit next_char = content[content.find(i) + len(i)] # check if the next character is a space if next_char == ' ': # replace the space with a non-breaking space content = content.replace(i + ' ', i + '\u00A0')
补充扩展需求
希望实现更复杂的替换规则:
- 句子
"Can the search be hypothetical?"中,将hypothetical与?之间的空格替换为非断空格; - 句子
"In the search it is necessary to refer to the "{figure 1.12}"中,将{与figure、figure与1.12、1.12与}之间的所有空格替换为非断空格。
尝试编写的代码存在语法和逻辑问题:
units = ['km', 'm', 'cm', 'mm', 'mi', 'yd', 'ft', 'in'] units_before_after = ['{'] nbsp = '\u00A0' rgx = re.sub(r'(\b\d+)(%s) (%s)\b'%(units, units_before_after),r'\1%s\2'%nbsp,text)) print(rgx)
解决方案
所有需求可通过正则表达式精准匹配目标位置实现,避免全局替换的误操作。以下是分场景的实现代码:
1. 数字与单位间的非断空格替换
匹配数字(支持小数)+空格+单位的结构,仅替换两者间的空格:
import os import re units = ['km', 'm', 'cm', 'mm', 'mi', 'yd', 'ft', 'in'] nbsp = '\u00A0' # 生成正则模式:匹配数字+空格+单位,单位作为整词匹配 unit_pattern = re.compile(r'(\d+\.?\d*) (%s)\b' % '|'.join(units)) # 遍历目录文件 for file in os.listdir(): if os.path.isfile(file): with open(file, 'r', encoding='utf-8') as f: content = f.read() # 替换匹配到的空格为非断空格 content = unit_pattern.sub(r'\1%s\2' % nbsp, content) # 写回文件(如需备份原文件,可先复制再覆盖) with open(file, 'w', encoding='utf-8') as f: f.write(content)
2. 单词与问号间的非断空格替换
匹配字母单词+空格+问号的结构,替换中间空格:
question_pattern = re.compile(r'([a-zA-Z]+) \?') content = question_pattern.sub(r'\1%s?' % nbsp, content)
3. 花括号内所有空格替换为非断空格
匹配花括号包裹的内容,将内部所有空格替换:
brace_pattern = re.compile(r'({[^}]+})') def replace_brace_spaces(match): return match.group(1).replace(' ', nbsp) content = brace_pattern.sub(replace_brace_spaces, content)
整合所有规则的完整代码
import os import re units = ['km', 'm', 'cm', 'mm', 'mi', 'yd', 'ft', 'in'] nbsp = '\u00A0' # 预编译所有正则模式 unit_pattern = re.compile(r'(\d+\.?\d*) (%s)\b' % '|'.join(units)) question_pattern = re.compile(r'([a-zA-Z]+) \?') brace_pattern = re.compile(r'({[^}]+})') def replace_brace_spaces(match): return match.group(1).replace(' ', nbsp) # 处理目录下所有文件 for file in os.listdir(): if os.path.isfile(file): with open(file, 'r', encoding='utf-8') as f: content = f.read() # 依次应用所有替换规则 content = unit_pattern.sub(r'\1%s\2' % nbsp, content) content = question_pattern.sub(r'\1%s?' % nbsp, content) content = brace_pattern.sub(replace_brace_spaces, content) with open(file, 'w', encoding='utf-8') as f: f.write(content)
内容的提问来源于stack exchange,提问作者Satanas
相关产品推荐
相关产品推荐

