You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何读取TXT文件中与日期关联的指定段落内容?

提取日期开头的完整段落解决方案

需求说明

需要从TXT文件中提取以日期(格式为「月份 年份.」)开头的段落,每个段落包含从该日期行开始到下一个日期行出现前的所有内容,同一日期的多行内容需合并到一起。

现有问题

当前代码仅能识别并输出日期所在行,无法捕获后续关联的非日期行内容。

解决方案代码

# 定义所有英文月份,确保覆盖所有可能的日期起始标识
months = ["January", "February", "March", "April", "May", "June", 
          "July", "August", "September", "October", "November", "December"]

# 用字典存储日期与对应段落内容
date_paragraphs = {}
current_date = None

with open("test.txt", encoding="utf8") as input_file:
    for line in input_file:
        stripped_line = line.strip()
        # 跳过空行(可根据实际需求删除此判断)
        if not stripped_line:
            continue
        
        # 判断当前行是否为日期起始行
        is_date_line = False
        for month in months:
            if stripped_line.startswith(f"{month} "):
                # 提取日期(截取到句号前的部分)
                current_date = stripped_line.split('.')[0].strip()
                # 初始化或追加当前日期的内容列表
                if current_date not in date_paragraphs:
                    date_paragraphs[current_date] = []
                date_paragraphs[current_date].append(stripped_line)
                is_date_line = True
                break
        
        # 非日期行且存在活跃日期时,追加内容到当前日期段落
        if not is_date_line and current_date is not None:
            date_paragraphs[current_date].append(stripped_line)

# 输出整理后的结果
for date, content_lines in date_paragraphs.items():
    print(f"=== {date} ===")
    print('\n'.join(content_lines))
    print("\n")

代码说明

  1. 月份列表:覆盖所有英文月份,用于精准识别日期起始行。
  2. 字典存储:以日期为键,对应值为该日期段落的所有行内容列表,方便后续合并输出。
  3. 行处理逻辑:
    • 读取每行后先去除首尾空白,可选择跳过空行优化输出。
    • 识别到日期行时,更新当前活跃日期,并将该行加入对应字典条目。
    • 非日期行直接追加到当前活跃日期的内容列表中。
  4. 结果输出:遍历字典,将每个日期的所有行合并为完整段落输出。

内容的提问来源于stack exchange,提问作者Hugo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 16:48:36