You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python正则表达式提取文件中指定标记区间内的行?

提取特定标记行之间内容的Python解决方案

我来帮你搞定这个问题!要提取文件中位于那两个writev标记行之间的内容,有两种实用的Python实现方法,根据你的文件大小和偏好来选就行:

方法一:逐行遍历(内存友好,适合大文件)

这种方法会逐行读取文件,通过一个标记位来判断是否进入目标区间,避免一次性加载整个文件到内存,非常适合处理大日志文件。

def extract_between_markers(file_path):
    # 定义起始和结束的标记片段(选足够唯一的部分就行)
    start_marker = '{"befor " 17}, {"androidID "'
    end_marker = '{"After " 16}, {"37abc5afce16b6www03", 17}'
    in_target_section = False
    result_lines = []
    
    with open(file_path, 'r', encoding='utf-8') as f:
        for line in f:
            cleaned_line = line.strip()
            # 遇到起始标记,开启收集模式,跳过标记行本身
            if start_marker in cleaned_line:
                in_target_section = True
                continue
            # 遇到结束标记,关闭收集模式,跳过标记行本身
            if end_marker in cleaned_line:
                in_target_section = False
                continue
            # 如果在收集模式内,就把当前行加入结果
            if in_target_section:
                result_lines.append(cleaned_line)
    
    return result_lines

# 调用示例,替换成你的文件路径
target_content = extract_between_markers('your_log_file.txt')
for line in target_content:
    print(line)

方法二:正则表达式匹配(代码简洁,适合小文件)

如果你的文件不大,可以一次性读入整个内容,用正则表达式直接匹配两个标记行之间的所有内容。注意要转义正则中的特殊字符(比如[、]、(、)),并且启用跨行匹配模式。

import re

def extract_with_regex(file_path):
    # 精确匹配起始和结束的标记行(注意转义特殊字符)
    start_line_pattern = r'\[pid \d+\] \d+:\d+:\d+ writev\(4, \[\{" hZZ v;", 11\}, \{" ", 1\}, \{"befor " 17\}, \{"androidID ", 10\}\], 4\) = 39\n'
    end_line_pattern = r'\n\[pid \d+\] \d+:\d+:\d+ writev\(3, \[\{"l ", 7\}, \{"hZZ ;", 11\}, \{" ", 1\}, \{"After " 16\}, \{"37abc5afce16b6www03", 17\}\], 5\)= 52'
    
    # 组合正则,用非贪婪模式匹配中间的所有内容,re.DOTALL让.匹配换行符
    full_pattern = re.compile(f'{start_line_pattern}(.*?){end_line_pattern}', re.DOTALL)
    
    with open(file_path, 'r', encoding='utf-8') as f:
        file_content = f.read()
    
    match_result = full_pattern.search(file_content)
    if match_result:
        # 把匹配到的内容按行拆分,过滤空行
        result_lines = [line.strip() for line in match_result.group(1).split('\n') if line.strip()]
        return result_lines
    else:
        print("未找到匹配的内容")
        return []

# 调用示例
target_content = extract_with_regex('your_log_file.txt')
for line in target_content:
    print(line)

两种方法的对比

  • 逐行遍历:内存占用低,适合GB级的大文件,逻辑直观好调试。
  • 正则匹配:代码更紧凑,适合小文件,但如果文件过大,一次性读入会占用较多内存。

内容的提问来源于stack exchange,提问作者kaloon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 03:40:05