You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PDFClown中如何覆盖子串关键词高亮,优先高亮完整关键词

Solution for Prioritizing Full Keyword Highlighting in PDFClown

Hey there, let's tackle this substring keyword highlighting issue in PDFClown. The core goal is to make sure longer, complete keywords (like just.ETS or Test.ETS) get highlighted instead of their shorter substrings (like ETS), right? Here's a step-by-step approach to make this work smoothly:

1. Preprocess Keywords to Prioritize Longer Terms

First, we need to sort our keyword list by length in descending order. This ensures that longer, complete terms are checked first, so they don't get overshadowed by their shorter substrings.

import java.util.Arrays;
import java.util.List;

List<String> keywords = Arrays.asList("ETS", "just.ETS", "Test.ETS");
// Sort keywords from longest to shortest to prioritize full terms
keywords.sort((a, b) -> Integer.compare(b.length(), a.length()));

2. Modify Highlighting Logic to Avoid Substring Overlaps

Next, we'll adjust the PDF text extraction and annotation logic in PDFClown. The key here is to track already highlighted regions so we don't re-highlight substrings within a fully matched keyword. We'll work with both the source ActualPdf and the output result PDF.

import org.pdfclown.documents.Page;
import org.pdfclown.documents.interaction.annotations.HighlightAnnotation;
import org.pdfclown.files.PdfDocument;
import org.pdfclown.tools.TextExtractor;
import org.pdfclown.util.math.Rectangle2D;

import java.io.FileInputStream;
import java.io.FileOutputStream;
import java.util.HashSet;
import java.util.Map;
import java.util.Set;

// Load source ActualPdf
PdfDocument actualPdf = new PdfDocument(new FileInputStream("path/to/your/actual.pdf"));
// Initialize result PDF for highlighted output
PdfDocument resultPdf = new PdfDocument();

// Iterate through each page in the source PDF
for (Page sourcePage : actualPdf.getPages()) {
    // Clone source page to result PDF
    Page resultPage = resultPdf.getPages().add(sourcePage.clone());
    
    TextExtractor textExtractor = new TextExtractor();
    Map<Rectangle2D, List<TextExtractor.TextChunk>> textChunks = textExtractor.extract(sourcePage);
    
    // Track regions already highlighted to prevent substring overlaps
    Set<Rectangle2D> highlightedRegions = new HashSet<>();
    
    for (String keyword : keywords) {
        for (Map.Entry<Rectangle2D, List<TextExtractor.TextChunk>> entry : textChunks.entrySet()) {
            Rectangle2D chunkBounds = entry.getKey();
            List<TextExtractor.TextChunk> chunks = entry.getValue();
            String chunkText = TextExtractor.TextChunk.toString(chunks);
            
            if (chunkText.contains(keyword)) {
                // Calculate exact bounds of the keyword within the text chunk
                Rectangle2D keywordBounds = calculateKeywordBounds(chunkBounds, chunks, keyword);
                
                // Skip if this region is already covered by a longer keyword
                boolean isOverlapped = highlightedRegions.stream()
                        .anyMatch(region -> region.intersects(keywordBounds));
                
                if (!isOverlapped) {
                    // Add yellow highlight annotation to result page
                    HighlightAnnotation highlight = new HighlightAnnotation(resultPage, keywordBounds);
                    highlight.setColor(new org.pdfclown.documents.contents.colorSpaces.DeviceRGBColor(1f, 1f, 0f));
                    resultPage.getAnnotations().add(highlight);
                    
                    // Add popup with associated measurement value
                    addMeasurementPopup(resultPage, keywordBounds, getMeasurementValue(keyword));
                    
                    // Mark this region as highlighted
                    highlightedRegions.add(keywordBounds);
                }
            }
        }
    }
}

// Save the final highlighted PDF
resultPdf.save(new FileOutputStream("path/to/your/highlighted_result.pdf"), org.pdfclown.files.SerializationModeEnum.STANDARD);

Helper Methods (You'll Need These)

  • calculateKeywordBounds: Computes the exact rectangular area of the keyword within the text chunk, using character positions and widths.
  • addMeasurementPopup: Creates a popup annotation attached to the highlight, displaying the associated measurement value for the keyword.
  • getMeasurementValue: Retrieves the measurement from a pre-defined mapping table (see step 4).

3. Configure PDFClown Jar Dependencies

Make sure your project has the correct PDFClown libraries set up:

  • If using Maven, add this to your pom.xml:
<dependency>
    <groupId>org.pdfclown</groupId>
    <artifactId>pdfclown-lib</artifactId>
    <version>0.1.2.3</version>
</dependency>
  • For manual setup, download the pdfclown-lib.jar and add it to your project's lib directory, then configure your IDE to include it as a library dependency.

4. Use a Mapping Table for Keyword-Measurement Pairs

To link full keywords to their corresponding popup measurement values, maintain a hash map. This ensures we only associate measurements with valid, full keywords (not substrings):

import java.util.HashMap;
import java.util.Map;

Map<String, String> keywordMeasurementMap = new HashMap<>();
keywordMeasurementMap.put("just.ETS", "15.2 cm");
keywordMeasurementMap.put("Test.ETS", "9.7 mm");
// Note: We don't include "ETS" here since it's a substring and shouldn't be highlighted

// Inside getMeasurementValue method:
private String getMeasurementValue(String keyword) {
    return keywordMeasurementMap.getOrDefault(keyword, "No measurement available");
}

内容的提问来源于stack exchange,提问作者Seshadri

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 06:17:50