基于正则表达式处理.msg文件[sel]块的内容修改及代码调试
需求说明
需处理.msg文件片段,完成以下操作:
- 定位所有
[sel xxx xxx]开头的块,处理块内的行 - 删除所有行中的
(+x)格式内容 - 在
[sel]块中,找出(+x)数值最大的行(多行数值相同则全部选中),在该行的初始标签(比如[s]或[s][f ...])之后、文本内容之前插入[clr 4]
输入示例
[msg MSG_015_0_0 [Makoto]] [s][bup 0 6 0 65535 1][f 4 10 65535 0 0][vp 8 0 0 0 65535 0]Dummy[n][f 1 3 65535][w][e] [sel SEL_016_0_0 top] [s]A(+3)[e] [s]B[e] [s]C(+3)[e] [msg MSG_021_0_0 [Makoto]] [s][bup 0 6 0 65535 1][f 4 10 65535 0 0][vp 8 0 0 0 65535 0]Sorry I had to ask something so[n]unreasonable... I appreciate you[n]going along with it, though.[n][f 1 3 65535][w][e] [sel SEL_023_0_0 top] [s]Fist! Of! Justice![e] [s]Report him to the police.[e] [s][f 0 8 2 1 2217]Try to reach out to her.(+3)[clr 27][e]
预期输出
[msg MSG_015_0_0 [Makoto]] [s][bup 0 6 0 65535 1][f 4 10 65535 0 0][vp 8 0 0 0 65535 0]Dummy[n][f 1 3 65535][w][e] [sel SEL_016_0_0 top] [s][clr 4]A[e] [s]B[e] [s][clr 4]C[e] [msg MSG_021_0_0 [Makoto]] [s][bup 0 6 0 65535 1][f 4 10 65535 0 0][vp 8 0 0 0 65535 0]Sorry I had to ask something so[n]unreasonable... I appreciate you[n]going along with it, though.[n][f 1 3 65535][w][e] [sel SEL_023_0_0 top] [s]Fist! Of! Justice![e] [s]Report him to the police.[e] [s][f 0 8 2 1 2217][clr 4]Try to reach out to her.[clr 27][e]
现有代码问题分析
现有代码存在以下关键问题:
- 按块整体处理而非逐行分析,无法给多行添加
[clr 4]标记 - 先替换
(+x)再查找原标记位置,此时原字符串已被修改,查找结果无效 - 使用
[col 4]而非需求中的[clr 4] - 插入位置逻辑错误,未准确定位到初始标签后、文本前的位置
修复后的代码
import re def process_sel_block(lines): # 收集每行的(+x)数值和原始内容 line_info = [] max_x = 0 for line in lines: match = re.search(r'\(\+(\d+)\)', line) x_val = int(match.group(1)) if match else 0 line_info.append((x_val, line)) if x_val > max_x: max_x = x_val processed_lines = [] for x_val, line in line_info: # 移除当前行的(+x)格式内容 cleaned_line = re.sub(r'\(\+\d+\)', '', line) # 给符合条件的行添加[clr 4] if x_val == max_x and max_x > 0: # 匹配[s]开头的所有连续标签部分 tag_end_match = re.match(r'(\[s\](?:\[[^\]]+\])*)', cleaned_line) if tag_end_match: tag_part = tag_end_match.group(1) rest_content = cleaned_line[len(tag_part):] cleaned_line = f"{tag_part}[clr 4]{rest_content}" processed_lines.append(cleaned_line) return processed_lines def main(): with open('input.msg', 'r', encoding='utf-8') as infile: content = infile.readlines() result = [] in_sel_block = False sel_block_lines = [] for line in content: stripped_line = line.strip() # 检测进入[sel]块 if stripped_line.startswith('[sel'): in_sel_block = True result.append(line) continue # 检测离开[sel]块 if in_sel_block: if stripped_line == '' or (not stripped_line.startswith('[s]') and not stripped_line.startswith('[e]')): processed_lines = process_sel_block(sel_block_lines) result.extend(processed_lines) result.append(line) in_sel_block = False sel_block_lines = [] else: sel_block_lines.append(line) else: result.append(line) # 处理文件末尾的[sel]块 if in_sel_block: processed_lines = process_sel_block(sel_block_lines) result.extend(processed_lines) with open('output.msg', 'w', encoding='utf-8') as outfile: outfile.writelines(result) if __name__ == '__main__': main()
代码说明
- 逐行识别
[sel]块:通过逐行读取文件,精准识别进入和离开[sel]块的时机,收集块内行进行针对性处理 - 提前记录数值信息:先遍历块内所有行,提取每行的
(+x)数值并找出最大值,避免修改字符串后丢失位置信息 - 精准插入标记:使用正则匹配
[s]开头的所有连续标签,在标签结束后、文本内容前插入[clr 4] - 先清理再标记:先移除所有
(+x)格式内容,再给符合条件的行添加标记,确保逻辑顺序正确
内容的提问来源于stack exchange,提问作者Medusa
相关产品推荐
相关产品推荐

