You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何筛选/tmp/file中job_type为BOX的完整条目并输出

解决方案

核心思路是缓存对应start_times行,当遇到insert行时判断job_type,仅保留BOX类型的条目及其对应的start_times行。

完整代码

with open('/tmp/file', 'r') as f:
    cached_start_line = None
    for raw_line in f:
        line = raw_line.strip()
        # 跳过空行
        if not line:
            continue
        
        # 缓存start_times行
        if line.startswith('start_times'):
            cached_start_line = line
        # 处理insert行
        elif line.startswith('insert'):
            # 检查是否为BOX类型
            if "job_type='BOX'" in line:
                if cached_start_line:
                    print(cached_start_line)
                    print(line)
            # 无论是否匹配,都重置缓存,准备下一组
            cached_start_line = None

代码逻辑说明

  1. 缓存机制:用cached_start_line变量临时存储当前insert行对应的start_times内容,确保成对输出。
  2. 行判断:
    • 遇到start_times开头的行,直接存入缓存;
    • 遇到insert行时,检查是否包含job_type='BOX':
      • 符合条件则先输出缓存的start_times行,再输出当前insert行;
      • 若为CMD类型,直接跳过,不输出任何内容。
  3. 重置缓存:处理完每一行insert后,强制重置缓存,避免残留上一组的内容干扰下一组判断。

适配调整

  • 如果文件中job_type使用双引号(如job_type="BOX"),只需把判断条件改为"job_type=\"BOX\"" in line;
  • 若需要将结果写入新文件而非打印,把print()替换为文件写入操作即可。

内容的提问来源于stack exchange,提问作者jada

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 20:02:32