如何使用Python从TXT文档中提取多组指定起止章节之间的文本内容
问题分析
- 匹配逻辑错误:原代码直接将单行去除空白后的内容与
start/end两个列表做相等判断,字符串和列表不可能相等,所以匹配逻辑完全失效,永远无法触发复制开关。 - 标签匹配规则不合理:示例文档中的章节标题并非完全等于列表内的标签文本(比如实际行内容为
Complaint / History Of Present Illness,标签仅为Complaint),用全等判断也无法匹配到对应行。 - 冗余代码:使用
with语法管理文件句柄时,代码块执行结束后会自动关闭文件,末尾手动调用outfile.close()属于多余操作。
修正后代码
intakedoc = 'Minnie.txt' savedfile = "Mouse.txt" start = ["Complaint", "Health History", "Psychiatric History and Treatment", "Education", "Drug and Alcohol Use", "Additional Personal Information", "Clinical Review of Systems", "Mental Status Examination" ] end = ["Health History", "Psychiatric History and Treatment", "Education", "Drug and Alcohol Use", "Additional Personal Information", "Clinical Review of Systems", "Mental Status Examination", "Risk Assessments"] with open(intakedoc, 'r', encoding='utf-8') as infile, open(savedfile, 'w', encoding='utf-8') as outfile: copy = False for line in infile: stripped_line = line.strip() # 判断当前行是否包含任意起始标签 if any(tag in stripped_line for tag in start): copy = True outfile.write(line) # 判断当前行是否包含任意结束标签 elif any(tag in stripped_line for tag in end): copy = False outfile.write(line) elif copy: outfile.write(line)
代码说明
- 调整了匹配逻辑:用
any()遍历start/end列表,判断当前行是否包含对应标签,适配章节标题带额外后缀的场景,如果你需要更精准的匹配,可以把tag in stripped_line改为stripped_line.startswith(tag),仅匹配以标签开头的行。 - 保留了原有写入规则:匹配到起始标签后开启复制,匹配到结束标签后关闭复制,起止标签行本身也会写入结果文件。
- 新增了
encoding='utf-8'参数,避免不同系统下读写文件出现乱码问题。 - 移除了冗余的手动关闭文件代码。
内容的提问来源于stack exchange,提问作者bmedclinic
相关产品推荐
相关产品推荐

