如何提取文件中两个指定标记字符串之间的特定范围文本?
文本提取实现方案
方案1:字符串分割法(最简单,不易出错,无需掌握正则规则)
直接用两个固定标识做两次分割即可,逻辑清晰不容易踩坑:
- 第一步:用前标识
Overall interfering behavior data trends are as followed:对原始文本做分割,取分割后的第二部分(即前标识之后的所有内容) - 第二步:对上一步得到的内容,用后标识
Observations of Client's response to skill acquisition:做分割,取分割后的第一部分,就是目标提取内容 - 最后对提取到的内容做首尾去空白处理即可
Python实现示例代码:
# 此处替换为你的原始文本内容 raw_text = """Observations of Client Behavior: Overall interfering behavior data trends are as followed: THIS IS THE DESIRED TEXT. Observations of Client's response to skill acquisition: Overall skill acquisition data trends ....""" # 定义前后分割标识 start_flag = "Overall interfering behavior data trends are as followed:" end_flag = "Observations of Client's response to skill acquisition:" # 两次分割提取目标内容 target_content = raw_text.split(start_flag)[1].split(end_flag)[0].strip() # 写入独立文本文件 with open("提取的行为观察内容.txt", "w", encoding="utf-8") as f: f.write(target_content)
方案2:正则匹配法(适合批量处理多段匹配的场景)
之前正则匹配失败大概率是没有开启单行匹配模式,默认正则的.不匹配换行符,如果文本包含换行就会匹配中断。正确的正则规则需要开启DOTALL模式,匹配规则为(?<=Overall interfering behavior data trends are as followed:).*?(?=Observations of Client's response to skill acquisition:)
(?<=xxx)是正向后顾断言,匹配xxx后面的位置.*?是非贪婪匹配所有字符,避免多段匹配时超出范围(?=xxx)是正向前瞻断言,匹配xxx前面的位置
Python实现示例代码:
import re # 此处替换为你的原始文本内容 raw_text = """Observations of Client Behavior: Overall interfering behavior data trends are as followed: THIS IS THE DESIRED TEXT. Observations of Client's response to skill acquisition: Overall skill acquisition data trends ....""" pattern = r'(?<=Overall interfering behavior data trends are as followed:).*?(?=Observations of Client\'s response to skill acquisition:)' # 开启DOTALL模式,让.可以匹配换行符 match_res = re.search(pattern, raw_text, re.DOTALL) if match_res: target_content = match_res.group().strip() # 写入独立文本文件 with open("提取的行为观察内容.txt", "w", encoding="utf-8") as f: f.write(target_content)
补充提示:如果待处理的是docx格式文档,可以先用
python-docx库把文档内容读取为字符串,再按上面的逻辑处理即可。
内容的提问来源于stack exchange,提问作者SSerb1989
相关产品推荐
相关产品推荐

