如何读取TXT文件中与日期关联的指定段落内容?
提取日期开头的完整段落解决方案
需求说明
需要从TXT文件中提取以日期(格式为「月份 年份.」)开头的段落,每个段落包含从该日期行开始到下一个日期行出现前的所有内容,同一日期的多行内容需合并到一起。
现有问题
当前代码仅能识别并输出日期所在行,无法捕获后续关联的非日期行内容。
解决方案代码
# 定义所有英文月份,确保覆盖所有可能的日期起始标识 months = ["January", "February", "March", "April", "May", "June", "July", "August", "September", "October", "November", "December"] # 用字典存储日期与对应段落内容 date_paragraphs = {} current_date = None with open("test.txt", encoding="utf8") as input_file: for line in input_file: stripped_line = line.strip() # 跳过空行(可根据实际需求删除此判断) if not stripped_line: continue # 判断当前行是否为日期起始行 is_date_line = False for month in months: if stripped_line.startswith(f"{month} "): # 提取日期(截取到句号前的部分) current_date = stripped_line.split('.')[0].strip() # 初始化或追加当前日期的内容列表 if current_date not in date_paragraphs: date_paragraphs[current_date] = [] date_paragraphs[current_date].append(stripped_line) is_date_line = True break # 非日期行且存在活跃日期时,追加内容到当前日期段落 if not is_date_line and current_date is not None: date_paragraphs[current_date].append(stripped_line) # 输出整理后的结果 for date, content_lines in date_paragraphs.items(): print(f"=== {date} ===") print('\n'.join(content_lines)) print("\n")
代码说明
- 月份列表:覆盖所有英文月份,用于精准识别日期起始行。
- 字典存储:以日期为键,对应值为该日期段落的所有行内容列表,方便后续合并输出。
- 行处理逻辑:
- 读取每行后先去除首尾空白,可选择跳过空行优化输出。
- 识别到日期行时,更新当前活跃日期,并将该行加入对应字典条目。
- 非日期行直接追加到当前活跃日期的内容列表中。
- 结果输出:遍历字典,将每个日期的所有行合并为完整段落输出。
内容的提问来源于stack exchange,提问作者Hugo
相关产品推荐
相关产品推荐

