You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提取TXT文件中包含"apple"的所有句子?已尝试正则等方法未果

解决提取含"apple"句子的问题

你的代码存在两个核心问题:

  • 未读取文件内容:re.findall() 需要传入字符串,直接传文件对象会报错,必须先读取文件内容。
  • 正则表达式局限性:原正则只匹配以.结尾的句子,忽略了!、?等常见句子结尾符号;且[^.]*不匹配换行符,会漏掉跨行的句子。

修正后的代码

import re

# 安全读取文件内容(with语句自动关闭文件)
with open("apple.txt", "r", encoding="utf-8") as fp:
    content = fp.read()

# 改进正则:匹配包含apple的句子,支持跨行、多种结尾标点
pattern = re.compile(r"([^.!?]*apple[^.!?]*[.!?])", re.DOTALL)
matched_sentences = pattern.findall(content)

# 输出结果(去除前后空白)
for sentence in matched_sentences:
    print(sentence.strip())

关键改进说明

  • 使用with语句操作文件,避免手动关闭文件可能引发的资源泄漏。
  • 添加re.DOTALL标志,让正则中的.匹配换行符,处理跨行的句子。
  • 将句子结尾标点扩展为.!?,覆盖英文中常见的句子结束符号。
  • 用strip()去除句子前后的空白字符,让输出更整洁。

进阶优化(处理缩写中的点)

如果文本中存在Mr.、U.S.A.这类带点的缩写,原正则会错误截断句子。可以用更精准的正则来规避:

# 排除常见缩写后的正则,仅匹配真正的句子结尾
pattern = re.compile(r"(?:(?<!Mr)(?<!Mrs)(?<!Ms)(?<!Dr)(?<!Sr)(?<!Jr)\.)|!|\?")
sentences = re.split(pattern, content)
matched_sentences = [s.strip() + p for s, p in zip(sentences, re.findall(pattern, content)) if "apple" in s]

for sentence in matched_sentences:
    print(sentence)

内容的提问来源于stack exchange,提问作者Chaitanya Chitturi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 08:35:30