You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在多份PDF文档中忽略大小写搜索指定关键词?

修改代码实现PDF关键词大小写不敏感搜索

你的代码要实现大小写不敏感匹配,核心是给正则搜索添加忽略大小写的标志,同时还要修复代码里的缩进错误。以下是修改后的完整代码:

# import packages
import PyPDF2
import re
import os, sys

dirs_list=[]
for root, dirs, files in os.walk(".", topdown=False):
    for name in dirs:
        dirs_list.append(dirs)

dirs_list=dirs_list[-1]
dirs_list.pop(0)
# print(dirs_list)

for k in dirs_list:
    # 修复缩进:这部分需放在循环内部
    data_names=os.listdir(k)
    data_names.pop(0)

    for j in data_names:
        # open the pdf file
        reader = PyPDF2.PdfReader(os.path.join(k, j))

        # define key terms
        strings = ['Lineare Gleichungen','Quadratische Gleichungen']
        total=len(strings)

        counter=0
        # extract text and do the search
        for page in reader.pages:
            text = page.extract_text()
            # 提前提取文本,避免重复调用提升效率
            if not text:
                continue  # 跳过无文本的页面
            for i in strings:
                # 添加re.IGNORECASE实现大小写不敏感匹配
                res_search = re.search(i, text, re.IGNORECASE)
                if res_search is not None:
                    counter+=1
    
    print(j, counter/total)

关键修改点说明:

  • 大小写不敏感匹配:在re.search(i, text, re.IGNORECASE)中加入re.IGNORECASE(可简写为re.I),正则引擎会忽略关键词和文本的大小写差异,比如Lineare Gleichungen能匹配lineare gleichungen、LINEARE GLEICHUNGEN等任意大小写组合。
  • 修复缩进错误:原代码中data_names=os.listdir(k)一行未缩进,导致语法错误,现在调整到for k in dirs_list:循环内部。
  • 优化文本提取:把text = page.extract_text()移到关键词循环外层,避免重复提取页面文本;同时增加判断,跳过未提取到文本的页面。

内容的提问来源于stack exchange,提问作者ghxk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 00:21:26