You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python从TXT文档中提取多组指定起止章节之间的文本内容

问题分析
  • 匹配逻辑错误:原代码直接将单行去除空白后的内容与start/end两个列表做相等判断,字符串和列表不可能相等,所以匹配逻辑完全失效,永远无法触发复制开关。
  • 标签匹配规则不合理:示例文档中的章节标题并非完全等于列表内的标签文本(比如实际行内容为Complaint / History Of Present Illness,标签仅为Complaint),用全等判断也无法匹配到对应行。
  • 冗余代码:使用with语法管理文件句柄时,代码块执行结束后会自动关闭文件,末尾手动调用outfile.close()属于多余操作。
修正后代码
intakedoc = 'Minnie.txt'
savedfile = "Mouse.txt"

start = ["Complaint", "Health History", "Psychiatric History and Treatment", "Education", "Drug and Alcohol Use", "Additional Personal Information", "Clinical Review of Systems", "Mental Status Examination" ]
end = ["Health History", "Psychiatric History and Treatment", "Education", "Drug and Alcohol Use", "Additional Personal Information", "Clinical Review of Systems", "Mental Status Examination", "Risk Assessments"]


with open(intakedoc, 'r', encoding='utf-8') as infile, open(savedfile, 'w', encoding='utf-8') as outfile:
    copy = False
    for line in infile:
        stripped_line = line.strip()
        # 判断当前行是否包含任意起始标签
        if any(tag in stripped_line for tag in start):
            copy = True
            outfile.write(line)
        # 判断当前行是否包含任意结束标签
        elif any(tag in stripped_line for tag in end):
            copy = False
            outfile.write(line)
        elif copy:
            outfile.write(line)
代码说明
  • 调整了匹配逻辑:用any()遍历start/end列表,判断当前行是否包含对应标签,适配章节标题带额外后缀的场景,如果你需要更精准的匹配,可以把tag in stripped_line改为stripped_line.startswith(tag),仅匹配以标签开头的行。
  • 保留了原有写入规则:匹配到起始标签后开启复制,匹配到结束标签后关闭复制,起止标签行本身也会写入结果文件。
  • 新增了encoding='utf-8'参数,避免不同系统下读写文件出现乱码问题。
  • 移除了冗余的手动关闭文件代码。

内容的提问来源于stack exchange,提问作者bmedclinic

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 13:54:02