如何在文本文件解析中实现带起始/终止短语的条件控制?
问题描述
我有两个列表start_phrases和stop_phrases,需要实现以下文件解析逻辑:
- 解析输入文件并写入输出文件时,当遇到仅包含
start_phrases中内容的行,开始将该行及后续连续行追加到输出文件 - 当遇到以
stop_phrases中内容开头的行,立即停止解析并终止循环,且不把该行写入输出文件
现有起始、终止短语定义:
start_phrases = ["Hello", "Come on:", "Introduction", "Background"] stop_phrases = ["This is provided to assist", "The background knowledge is to know"]
当前读取文件的代码如下:
with open (data, "r", encoding='utf-8') as myfile: for line in myfile: line.strip() print(line)
请问如何为这段代码添加上述解析条件?
解决方案
你可以通过添加状态标记控制写入逻辑,完整实现代码如下:
start_phrases = ["Hello", "Come on:", "Introduction", "Background"] stop_phrases = ["This is provided to assist", "The background knowledge is to know"] input_file = "your_input.txt" # 替换为实际输入文件路径 output_file = "your_output.txt" # 替换为实际输出文件路径 write_mode = False # 标记是否开启写入 # 同时打开输入和输出文件,自动处理关闭 with open(input_file, "r", encoding='utf-8') as infile, open(output_file, "a", encoding='utf-8') as outfile: for line in infile: stripped_line = line.strip() # 触发起始条件:行仅包含起始短语(去空白后完全匹配) if stripped_line in start_phrases: write_mode = True outfile.write(line) continue # 处于写入模式时,先检查终止条件 if write_mode: # 行以任意终止短语开头则停止循环 if any(line.startswith(phrase) for phrase in stop_phrases): break # 写入当前行 outfile.write(line)
关键逻辑说明
stripped_line = line.strip():strip()返回新字符串,必须重新赋值才能用去空白后的内容判断stripped_line in start_phrases:确保行仅包含起始短语(去除前后空白后完全匹配)any(line.startswith(phrase) for phrase in stop_phrases):高效判断行是否以任意终止短语开头- 输出文件用
"a"模式实现追加写入,若需要覆盖原有内容可改为"w"模式 with语句同时管理输入输出文件,避免手动关闭文件导致的资源泄漏
内容的提问来源于stack exchange,提问作者Lilly
相关产品推荐
相关产品推荐

