You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于正则表达式处理.msg文件[sel]块的内容修改及代码调试

需求说明

需处理.msg文件片段,完成以下操作:

  1. 定位所有[sel xxx xxx]开头的块,处理块内的行
  2. 删除所有行中的(+x)格式内容
  3. 在[sel]块中,找出(+x)数值最大的行(多行数值相同则全部选中),在该行的初始标签(比如[s]或[s][f ...])之后、文本内容之前插入[clr 4]

输入示例

[msg MSG_015_0_0 [Makoto]]
[s][bup 0 6 0 65535 1][f 4 10 65535 0 0][vp 8 0 0 0 65535 0]Dummy[n][f 1 3 65535][w][e]

[sel SEL_016_0_0 top]
[s]A(+3)[e]
[s]B[e]
[s]C(+3)[e]


[msg MSG_021_0_0 [Makoto]]
[s][bup 0 6 0 65535 1][f 4 10 65535 0 0][vp 8 0 0 0 65535 0]Sorry I had to ask something so[n]unreasonable... I appreciate you[n]going along with it, though.[n][f 1 3 65535][w][e]

[sel SEL_023_0_0 top]
[s]Fist! Of! Justice![e]
[s]Report him to the police.[e]
[s][f 0 8 2 1 2217]Try to reach out to her.(+3)[clr 27][e]

预期输出

[msg MSG_015_0_0 [Makoto]]
[s][bup 0 6 0 65535 1][f 4 10 65535 0 0][vp 8 0 0 0 65535 0]Dummy[n][f 1 3 65535][w][e]

[sel SEL_016_0_0 top]
[s][clr 4]A[e]
[s]B[e]
[s][clr 4]C[e]


[msg MSG_021_0_0 [Makoto]]
[s][bup 0 6 0 65535 1][f 4 10 65535 0 0][vp 8 0 0 0 65535 0]Sorry I had to ask something so[n]unreasonable... I appreciate you[n]going along with it, though.[n][f 1 3 65535][w][e]

[sel SEL_023_0_0 top]
[s]Fist! Of! Justice![e]
[s]Report him to the police.[e]
[s][f 0 8 2 1 2217][clr 4]Try to reach out to her.[clr 27][e]

现有代码问题分析

现有代码存在以下关键问题:

  • 按块整体处理而非逐行分析,无法给多行添加[clr 4]标记
  • 先替换(+x)再查找原标记位置,此时原字符串已被修改,查找结果无效
  • 使用[col 4]而非需求中的[clr 4]
  • 插入位置逻辑错误,未准确定位到初始标签后、文本前的位置

修复后的代码

import re

def process_sel_block(lines):
    # 收集每行的(+x)数值和原始内容
    line_info = []
    max_x = 0
    for line in lines:
        match = re.search(r'\(\+(\d+)\)', line)
        x_val = int(match.group(1)) if match else 0
        line_info.append((x_val, line))
        if x_val > max_x:
            max_x = x_val
    
    processed_lines = []
    for x_val, line in line_info:
        # 移除当前行的(+x)格式内容
        cleaned_line = re.sub(r'\(\+\d+\)', '', line)
        # 给符合条件的行添加[clr 4]
        if x_val == max_x and max_x > 0:
            # 匹配[s]开头的所有连续标签部分
            tag_end_match = re.match(r'(\[s\](?:\[[^\]]+\])*)', cleaned_line)
            if tag_end_match:
                tag_part = tag_end_match.group(1)
                rest_content = cleaned_line[len(tag_part):]
                cleaned_line = f"{tag_part}[clr 4]{rest_content}"
        processed_lines.append(cleaned_line)
    return processed_lines

def main():
    with open('input.msg', 'r', encoding='utf-8') as infile:
        content = infile.readlines()
    
    result = []
    in_sel_block = False
    sel_block_lines = []
    
    for line in content:
        stripped_line = line.strip()
        # 检测进入[sel]块
        if stripped_line.startswith('[sel'):
            in_sel_block = True
            result.append(line)
            continue
        # 检测离开[sel]块
        if in_sel_block:
            if stripped_line == '' or (not stripped_line.startswith('[s]') and not stripped_line.startswith('[e]')):
                processed_lines = process_sel_block(sel_block_lines)
                result.extend(processed_lines)
                result.append(line)
                in_sel_block = False
                sel_block_lines = []
            else:
                sel_block_lines.append(line)
        else:
            result.append(line)
    
    # 处理文件末尾的[sel]块
    if in_sel_block:
        processed_lines = process_sel_block(sel_block_lines)
        result.extend(processed_lines)
    
    with open('output.msg', 'w', encoding='utf-8') as outfile:
        outfile.writelines(result)

if __name__ == '__main__':
    main()

代码说明

  1. 逐行识别[sel]块:通过逐行读取文件,精准识别进入和离开[sel]块的时机,收集块内行进行针对性处理
  2. 提前记录数值信息:先遍历块内所有行,提取每行的(+x)数值并找出最大值,避免修改字符串后丢失位置信息
  3. 精准插入标记:使用正则匹配[s]开头的所有连续标签,在标签结束后、文本内容前插入[clr 4]
  4. 先清理再标记:先移除所有(+x)格式内容,再给符合条件的行添加标记,确保逻辑顺序正确

内容的提问来源于stack exchange,提问作者Medusa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 10:16:06