You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在批量PDF中实现多关键词搜索?Python工具优化求助

PDF多关键词搜索GUI工具优化方案

原工具仅支持单关键词搜索,需优化为支持多关键词(支持任意匹配或全匹配两种模式),同时精简冗余代码、提升用户体验。

核心优化要点

  • 多关键词解析:将输入内容按逗号分割,自动去除关键词前后空格,生成有效关键词列表
  • 搜索逻辑升级:提供两种匹配模式,可按需切换
    • 任意匹配:文档包含至少一个关键词即命中
    • 全匹配:文档必须包含所有输入关键词才命中
  • 代码精简:移除未使用的库导入(requests、webbrowser、BeautifulSoup),删除重复的import os
  • 路径安全处理:用os.path.join()拼接文件路径,避免跨平台路径格式问题
  • GUI体验优化:修正输入提示文案,新增结果展示文本框,让用户直接在GUI内查看搜索结果

修改后的完整代码

import os
import fitz
import customtkinter

# 目标文件夹路径
path = r'O:\Sent Questions'
files = os.listdir(path)

# 初始化GUI
customtkinter.set_appearance_mode("dark")
root = customtkinter.CTk()
root.geometry("700x500")  # 扩大窗口容纳结果展示区
root.title("Questions Keyword Search")

# 标题标签
label = customtkinter.CTkLabel(root, text="Questions Keyword Search Engine", font=("Inter", 30))
label.pack(side=TOP, pady=20)

# 结果展示文本框
result_text = customtkinter.CTkTextbox(root, width=650, height=200)
result_text.pack(pady=10)

def search():
    # 获取并解析输入的关键词
    input_str = entry.get().strip()
    if not input_str:
        result_text.delete(1.0, END)
        result_text.insert(END, "请输入关键词,用逗号分隔")
        return
    
    # 分割并清洗关键词
    keywords = [kw.strip() for kw in input_str.split(',') if kw.strip()]
    result_text.delete(1.0, END)
    matched_files = []

    for file in files:
        # 仅处理PDF文件
        if not file.lower().endswith('.pdf'):
            continue
        
        file_path = os.path.join(path, file)
        try:
            doc = fitz.open(file_path)
            file_matched = False
            for page in doc:
                text = page.get_text()
                # --- 切换匹配模式 ---
                # 模式1:任意关键词匹配(满足一个即命中)
                # if any(kw in text for kw in keywords):
                #     file_matched = True
                #     break
                
                # 模式2:所有关键词全匹配(必须全部包含)
                if all(kw in text for kw in keywords):
                    file_matched = True
                    break
            
            if file_matched:
                matched_files.append(file)
        except Exception as e:
            result_text.insert(END, f"处理文件{file}时出错:{str(e)}\n")
    
    # 展示搜索结果
    if matched_files:
        result_text.insert(END, "找到匹配的文件:\n")
        for idx, f in enumerate(matched_files, 1):
            result_text.insert(END, f"{idx}. {f}\n")
    else:
        result_text.insert(END, "未找到匹配的文件")

# 输入提示区域
label_1 = customtkinter.CTkLabel(root, text="Enter Keywords Below", font=("Inter", 15))
label_1.pack(pady=(10,0))

label_2 = customtkinter.CTkLabel(root, text="输入多个关键词请用逗号分隔", font=("Inter", 12))
label_2.pack()

label_3 = customtkinter.CTkLabel(root, text='Example: cocaine, SoHT, LC-MS', font=("Inter", 12))
label_3.pack(pady=(0,10))

# 输入框与搜索按钮
entry = customtkinter.CTkEntry(master=root, width=300)
entry.pack(pady=5)
button = customtkinter.CTkButton(master=root, text="Search", command=search)
button.pack(pady=10)

root.mainloop()

关键说明

  • 代码默认启用全匹配模式,若需切换为任意匹配,注释掉all()相关代码,取消any()代码的注释即可
  • 新增PDF文件过滤逻辑,避免处理非PDF格式的文件
  • 添加异常捕获,处理文件打开失败等异常情况
  • 结果直接展示在GUI文本框中,无需切换到控制台查看

内容的提问来源于stack exchange,提问作者Lewis Simcoates

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 08:35:28