You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python统计PDF文本片段中给定列表的单词出现次数

问题原因及修复方案
  • 计数变量被重复重置:你在遍历每一页的循环里写了for elem in words: count[elem] = 0,每处理一页就会把之前所有页的计数清零,最终只会保留最后一页的统计结果,若最后一页没有匹配关键词就会全部返回0。
  • 仅提取了PDF的表格内容,未处理正文文本:代码中只调用了extract_tables()方法,只能获取PDF中的表格数据,正文段落的内容完全没有纳入统计范围。
  • 表格内容遍历逻辑错误:for line in f'{i} --- {tbl}'是对拼接后的字符串做遍历,每次循环拿到的line是单个字符,后续split()操作完全无意义,根本没拿到表格内的实际文本。
  • 关键词带多余后缀空格:你的关键词列表里warrant 、combination 两个词末尾带了空格,只有当PDF中该词后恰好跟空格时才能匹配,若跟标点、换行等就会匹配失败。

修正后的代码

import pdfplumber

pdf_file = "CapitalCorp.pdf"
# 去掉关键词末尾多余空格
words = ['blank','warrant','offering','combination','SPAC','founders']
# 计数初始化放在循环外,只执行一次
count = {elem:0 for elem in words}

with pdfplumber.open(pdf_file) as pdf:
    for pg in pdf.pages:
        # 先统计当前页普通正文文本的关键词
        page_text = pg.extract_text()
        if page_text:
            elements = page_text.split()
            for word in words:
                count[word] += elements.count(word)
        # 再统计当前页表格内的关键词
        tables = pg.extract_tables()
        for table in tables:
            for row in table:
                # 遍历表格每一行的所有单元格
                for cell in row:
                    if cell:
                        cell_elements = str(cell).split()
                        for word in words:
                            count[word] += cell_elements.count(word)
print(count)

如果需要适配大小写不同的匹配场景,可以把分割后的元素和关键词都转成小写再匹配,比如将elements.count(word)修改为[e.lower() for e in elements].count(word.lower())即可。

内容的提问来源于stack exchange,提问作者Math4264

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 12:24:03