在Notepad++中编写Python脚本实现匹配项按行递增编号替换
问题分析
原脚本的核心问题有两个:
- 计数变量
i在文件循环外初始化,仅按文件递增,而非文件内的匹配项; - 使用
str.replace()会一次性替换文件内所有匹配符号,导致所有结果共用同一个编号。
要实现「每个文件内逐匹配项递增编号,且每个文件重置计数」,需要改用支持逐匹配回调的替换方式,或逐行处理并逐个替换。
解决方案1:正则回调替换(简洁高效,适合多数场景)
利用re.sub()的回调函数特性,每次匹配时自动递增计数:
# -*- coding: utf-8 -*- import os import re # 预编译匹配模式(固定字符串也可用正则,方便后续扩展) search_pattern = re.compile('▶️ ') suffix = '.-' for root, dirs, files in os.walk('C:\\temp\\prueba'): for file in files: fname, ext = os.path.splitext(file) if ext == '.txt': full_path = os.path.join(root, file) # 每个文件初始化计数为0 count = 0 # 定义替换回调函数,每次匹配递增计数 def replace_func(match): nonlocal count count += 1 return f"{match.group()}{count}{suffix}" # 读取+替换+写入文件,指定utf-8编码避免特殊符号乱码 with open(full_path, 'r', encoding='utf-8') as f: content = f.read() updated_content = search_pattern.sub(replace_func, content) with open(full_path, 'w', encoding='utf-8') as f: f.write(updated_content)
解决方案2:逐行处理(适合超大文件,低内存占用)
如果处理的是大文件,逐行读取并逐个替换可避免一次性加载整个文件到内存:
# -*- coding: utf-8 -*- import os target_str = '▶️ ' suffix = '.-' for root, dirs, files in os.walk('C:\\temp\\prueba'): for file in files: fname, ext = os.path.splitext(file) if ext == '.txt': full_path = os.path.join(root, file) count = 0 updated_lines = [] with open(full_path, 'r', encoding='utf-8') as f: for line in f: # 统计当前行的匹配次数,逐个替换递增 match_count = line.count(target_str) if match_count > 0: temp_line = line for _ in range(match_count): count += 1 # 每次只替换1个匹配项 temp_line = temp_line.replace(target_str, f"{target_str}{count}{suffix}", 1) updated_lines.append(temp_line) else: updated_lines.append(line) with open(full_path, 'w', encoding='utf-8') as f: f.writelines(updated_lines)
关键改动说明
- 计数重置:将计数变量移到每个
.txt文件的处理逻辑内部,确保每个文件开始时计数归0; - 逐匹配替换:放弃一次性替换的
str.replace(),改用正则回调或逐行逐个替换,实现每次匹配递增编号; - 编码兼容:读写文件时显式指定
encoding='utf-8',避免特殊符号(如▶️)出现乱码问题。
内容的提问来源于stack exchange,提问作者Regue
相关产品推荐
相关产品推荐

