You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修改Python日志关键词搜索脚本以提取含关键词的句子?

日志文件关键词搜索:提取单个句子而非整行的解决方案

问题分析

原脚本能够定位关键词的位置,但当一行包含多个句子时,会输出整行内容。我们可以通过两种思路解决这个问题,同时优化原脚本的鲁棒性(比如自动管理文件句柄、处理一行多关键词的情况)。

方案1:提取关键词所在的完整句子(按句号分割)

核心逻辑是:将当前行按句号分割为多个句子,定位包含关键词的句子,同时保留原索引的准确性。

修改后的完整代码:

import os

# 输出文件路径
output_path = "D:\\X250\\Python_Scripts\\Search_File_for_Keyword_and_Print_Line\\Results.txt"

# 获取用户输入
search_path = input("Enter directory path to search : ")
file_type = input("File Type : ")
search_str = input("Enter the search string : ")

# 处理路径格式
if not (search_path.endswith("/") or search_path.endswith("\\")):
    search_path = search_path + "\\"
if not os.path.exists(search_path):
    search_path = "."

# 写入结果文件(使用with自动管理文件)
with open(output_path, 'w', encoding='utf-8') as fw:
    # 遍历目录下的文件
    for fname in os.listdir(search_path):
        if fname.endswith(file_type):
            file_full_path = os.path.join(search_path, fname)
            with open(file_full_path, 'r', encoding='utf-8') as fo:
                line_no = 1
                for line in fo:
                    line = line.rstrip('\n')  # 去掉换行符
                    start_idx = 0
                    # 循环查找当前行中所有关键词的位置
                    while True:
                        index = line.find(search_str, start_idx)
                        if index == -1:
                            break
                        # 找关键词所在句子的起始(前一个句号之后)
                        sentence_start = line.rfind('.', 0, index) + 1
                        # 找关键词所在句子的结束(后一个句号之前)
                        sentence_end = line.find('.', index + len(search_str))
                        if sentence_end == -1:
                            sentence_end = len(line)
                        # 提取句子(去除前后空格)
                        target_sentence = line[sentence_start:sentence_end].strip()
                        # 输出到控制台
                        print(f"{fname} [{line_no}, {index}] {target_sentence}")
                        # 写入结果文件
                        fw.write(f"{fname} {line_no} {index}  {target_sentence}\n")
                        # 更新起始索引,查找下一个关键词
                        start_idx = index + len(search_str)
                    line_no += 1

代码说明

  • 使用with语句自动管理文件句柄,避免手动关闭文件的遗漏
  • 用os.path.join拼接路径,适配不同操作系统的路径格式
  • 循环查找一行中的所有关键词,避免遗漏多个匹配的情况
  • 通过rfind和find定位关键词前后的句号,提取精准的句子内容

方案2:提取关键词前后固定长度的上下文

如果日志中的句子分隔不规范(比如没有用句号结尾),可以选择提取关键词前后固定长度的上下文,保证展示关键词的语境。

修改后的核心代码部分(替换方案1中的句子提取逻辑):

# 定义上下文长度
context_length = 50
# ... 其他代码不变 ...
while True:
    index = line.find(search_str, start_idx)
    if index == -1:
        break
    # 计算上下文的起始和结束位置,避免索引越界
    context_start = max(0, index - context_length)
    context_end = min(len(line), index + len(search_str) + context_length)
    # 提取上下文内容
    target_context = line[context_start:context_end].strip()
    # 给关键词添加星号标记,方便快速定位(可选)
    target_context = target_context.replace(search_str, f"*{search_str}*")
    # 输出和写入
    print(f"{fname} [{line_no}, {index}] {target_context}")
    fw.write(f"{fname} {line_no} {index}  {target_context}\n")
    start_idx = index + len(search_str)

代码说明

  • 通过context_length自定义上下文的长度
  • 使用max和min避免索引越界
  • 可选给关键词添加星号标记,提升可读性

原脚本的其他优化点

  • 增加了encoding='utf-8'参数,避免读取/写入文件时出现编码错误
  • 替换冗余的readline循环为for line in fo,代码更简洁
  • 处理了一行中多个关键词的情况,不会遗漏匹配

内容的提问来源于stack exchange,提问作者Col

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 20:09:59