You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PDF文件短语计数异常问题排查及正则修正方案咨询

PDF指定短语计数异常的解决

输入

  • 包含指定短语集合的PDF文件

需求

统计并输出PDF文件中每个指定短语的出现次数

问题

编写的代码对所有短语的计数始终返回1,推测是代码未遍历整个文件统计所有出现次数导致的。

原代码

import org.apache.pdfbox.Loader;
import org.apache.pdfbox.pdmodel.PDDocument;
import org.apache.pdfbox.text.PDFTextStripper;

import java.io.File;
import java.util.HashMap;
import java.util.Map;

public class task {
    public static void main(String[] args) {
        try {
            PDDocument document = Loader.loadPDF(new File("filepath"));
            PDFTextStripper textStripper = new PDFTextStripper();
            String pdfText = textStripper.getText(document);
            String[] sentencesToSearch = {
                    "hotels, restaurants, airlines, golf clubs, food and beverage, meeting, theme park, cruise, casino",
                    "revenue management techniques",
                    "Excel",
                    "data",
                    "marketing",
                    "finance",
                    "cost",
                    "economic principles",
                    "strategy",
                    "knowledge of",
                    "that"
            };
            Map<String, Integer> sentenceCount = new HashMap<>();

            for (String sentence : sentencesToSearch) {
                    if (pdfText.contains(sentence)) {
                        sentenceCount.put(sentence, sentenceCount.getOrDefault(sentence, 0) + 1);
                    }
            }
            document.close();
            for (Map.Entry<String, Integer> entry : sentenceCount.entrySet()) {
                System.out.println("Sentence: " + entry.getKey());
                System.out.println("Count: " + entry.getValue());
            }
        } catch (Exception e) {
            e.printStackTrace();
        }
    }
}

问题原因

原代码中pdfText.contains(sentence)仅判断短语是否在文本中存在,一旦存在就将计数加1,不会统计短语出现的总次数,因此每个短语的计数最多为1。

可行的正则实现代码

for (String word : definedWords) {                 
   Pattern pattern = Pattern.compile("\\b" + Pattern.quote(word) + "\\b");                 
   Matcher matcher = pattern.matcher(pdfText);                   
   int count = 0;                 
   while (matcher.find()) {                      
      count++;                 
   }                 
   phraseCount.put(word, count);              
}

代码说明

通过正则表达式的Matcher.find()方法循环遍历文本,每次找到匹配的短语就将计数加1,最终得到该短语在PDF文本中的总出现次数。Pattern.quote(word)用于处理短语中包含正则特殊字符的情况,\\b用于匹配单词边界,避免部分匹配(比如避免将"dat"统计到"data"中)。

内容的提问来源于stack exchange,提问作者Ramya Podha

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 15:11:10