You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python正则提取多行文件中符合条件的标记包裹内容块

实现方案

方案1:正则匹配实现

适合中小体积文件,代码简洁直接:

import re

# 读取源文件内容
with open("源文件路径.txt", "r", encoding="utf-8") as f:
    content = f.read()

# 正则匹配规则:匹配<Start开头、首行内容为1;、最近的<End结尾的块
pattern = re.compile(r"<Start\s*^1;.*?<End", re.DOTALL | re.MULTILINE)
matched_blocks = pattern.findall(content)

# 拼接结果并输出/写入文件
result = "\n".join(matched_blocks)
print(result)

with open("输出文件路径.txt", "w", encoding="utf-8") as f:
    f.write(result)

规则说明:

  • re.DOTALL:让通配符.可以匹配换行符,支持跨行匹配
  • re.MULTILINE:让^标识符可以匹配每一行的行首,确保校验的是块内第一行内容
  • 非贪婪量词.*?:保证每次匹配到最近的<End就停止,不会跨多个块吞内容

方案2:逐行状态机实现

适合超大体积文件,无需一次性加载全部内容到内存,逻辑更稳定不易踩正则匹配坑:

in_valid_block = False
current_block = []
result = []

with open("源文件路径.txt", "r", encoding="utf-8") as f:
    for line in f:
        stripped = line.strip()
        if stripped.startswith("<Start"):
            in_valid_block = True
            current_block = [line]
            is_first_line = True
            continue
        if stripped.startswith("<End"):
            if in_valid_block:
                current_block.append(line)
                result.extend(current_block)
            in_valid_block = False
            continue
        if in_valid_block:
            current_block.append(line)
            if is_first_line:
                # 校验块内第一行是否为1;
                if stripped != "1;":
                    in_valid_block = False
                    current_block = []
                is_first_line = False

# 输出/写入结果
final_content = "".join(result)
print(final_content)
with open("输出文件路径.txt", "w", encoding="utf-8") as f:
    f.write(final_content)

内容的提问来源于stack exchange,提问作者PythonNewbie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 13:24:03