You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

是否应为读取文件并打印含指定关键词行的函数添加正则表达式?

问题分析与修正

先说说你代码里的几个核心问题:

  • 函数参数filename完全没用到,反而硬编码打开了gatsby.txt,导致函数失去通用性
  • 你用keyword = input()覆盖了传入的参数值,如果要让用户输入关键词,应该去掉函数的keyword参数,或者不要在这里重新赋值
  • 循环变量用了keyword,和变量名冲突,且逻辑完全错误:遍历sentences却判断keyword in filename,根本不是“找包含关键词的句子”的逻辑
  • 用split('.')分割句子太粗糙,句子可能以!、?结尾,分割后会丢失标点,还会出现空字符串(比如连续句号的情况)

用正则实现精准匹配的修正版本

下面分两种常见需求给出实现:

需求1:匹配包含关键词的句子(忽略大小写,支持部分匹配)

比如找包含gatsby的句子,不管大小写,不管是单独单词还是单词的一部分:

import re

def print_lines_with_keyword(filename):
    keyword = input("请输入要查找的关键词: ")
    # 编译正则,re.IGNORECASE忽略大小写,re.escape转义特殊字符
    pattern = re.compile(re.escape(keyword), re.IGNORECASE)
    
    with open(filename, 'r', encoding='utf-8') as file:
        content = file.read()
        # 用正则分割句子,保留结尾标点,比split('.')更准确
        sentences = re.split(r'(?<=[.!?])\s+', content)
        
        for sentence in sentences:
            if pattern.search(sentence):
                print(sentence.strip())

需求2:匹配包含完整关键词的句子(精确匹配单词,忽略大小写)

比如找包含单独单词gatsby的句子,排除gatsbys这类变体:

import re

def print_lines_with_keyword(filename):
    keyword = input("请输入要查找的关键词: ")
    # \b匹配单词边界,确保是完整单词
    pattern = re.compile(r'\b' + re.escape(keyword) + r'\b', re.IGNORECASE)
    
    with open(filename, 'r', encoding='utf-8') as file:
        content = file.read()
        sentences = re.split(r'(?<=[.!?])\s+', content)
        
        for sentence in sentences:
            if pattern.search(sentence):
                print(sentence.strip())

关键细节解释

  • re.escape(keyword):转义关键词里的正则特殊字符(比如.、*),避免语法错误
  • re.IGNORECASE:让匹配忽略大小写,Gatsby和gatsby都能被匹配到
  • re.split(r'(?<=[.!?])\s+', content):用正向断言分割,确保分割点在.!?后的空格,分割出的句子会保留结尾标点
  • 循环遍历sentences,用pattern.search(sentence)检查句子是否包含匹配的关键词,找到就打印

内容的提问来源于stack exchange,提问作者floofcookie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 08:05:55