如何修改Python日志关键词搜索脚本以提取含关键词的句子?
日志文件关键词搜索:提取单个句子而非整行的解决方案
问题分析
原脚本能够定位关键词的位置,但当一行包含多个句子时,会输出整行内容。我们可以通过两种思路解决这个问题,同时优化原脚本的鲁棒性(比如自动管理文件句柄、处理一行多关键词的情况)。
方案1:提取关键词所在的完整句子(按句号分割)
核心逻辑是:将当前行按句号分割为多个句子,定位包含关键词的句子,同时保留原索引的准确性。
修改后的完整代码:
import os # 输出文件路径 output_path = "D:\\X250\\Python_Scripts\\Search_File_for_Keyword_and_Print_Line\\Results.txt" # 获取用户输入 search_path = input("Enter directory path to search : ") file_type = input("File Type : ") search_str = input("Enter the search string : ") # 处理路径格式 if not (search_path.endswith("/") or search_path.endswith("\\")): search_path = search_path + "\\" if not os.path.exists(search_path): search_path = "." # 写入结果文件(使用with自动管理文件) with open(output_path, 'w', encoding='utf-8') as fw: # 遍历目录下的文件 for fname in os.listdir(search_path): if fname.endswith(file_type): file_full_path = os.path.join(search_path, fname) with open(file_full_path, 'r', encoding='utf-8') as fo: line_no = 1 for line in fo: line = line.rstrip('\n') # 去掉换行符 start_idx = 0 # 循环查找当前行中所有关键词的位置 while True: index = line.find(search_str, start_idx) if index == -1: break # 找关键词所在句子的起始(前一个句号之后) sentence_start = line.rfind('.', 0, index) + 1 # 找关键词所在句子的结束(后一个句号之前) sentence_end = line.find('.', index + len(search_str)) if sentence_end == -1: sentence_end = len(line) # 提取句子(去除前后空格) target_sentence = line[sentence_start:sentence_end].strip() # 输出到控制台 print(f"{fname} [{line_no}, {index}] {target_sentence}") # 写入结果文件 fw.write(f"{fname} {line_no} {index} {target_sentence}\n") # 更新起始索引,查找下一个关键词 start_idx = index + len(search_str) line_no += 1
代码说明
- 使用
with语句自动管理文件句柄,避免手动关闭文件的遗漏 - 用
os.path.join拼接路径,适配不同操作系统的路径格式 - 循环查找一行中的所有关键词,避免遗漏多个匹配的情况
- 通过
rfind和find定位关键词前后的句号,提取精准的句子内容
方案2:提取关键词前后固定长度的上下文
如果日志中的句子分隔不规范(比如没有用句号结尾),可以选择提取关键词前后固定长度的上下文,保证展示关键词的语境。
修改后的核心代码部分(替换方案1中的句子提取逻辑):
# 定义上下文长度 context_length = 50 # ... 其他代码不变 ... while True: index = line.find(search_str, start_idx) if index == -1: break # 计算上下文的起始和结束位置,避免索引越界 context_start = max(0, index - context_length) context_end = min(len(line), index + len(search_str) + context_length) # 提取上下文内容 target_context = line[context_start:context_end].strip() # 给关键词添加星号标记,方便快速定位(可选) target_context = target_context.replace(search_str, f"*{search_str}*") # 输出和写入 print(f"{fname} [{line_no}, {index}] {target_context}") fw.write(f"{fname} {line_no} {index} {target_context}\n") start_idx = index + len(search_str)
代码说明
- 通过
context_length自定义上下文的长度 - 使用
max和min避免索引越界 - 可选给关键词添加星号标记,提升可读性
原脚本的其他优化点
- 增加了
encoding='utf-8'参数,避免读取/写入文件时出现编码错误 - 替换冗余的
readline循环为for line in fo,代码更简洁 - 处理了一行中多个关键词的情况,不会遗漏匹配
内容的提问来源于stack exchange,提问作者Col
相关产品推荐
相关产品推荐

