PDFClown中如何覆盖子串关键词高亮,优先高亮完整关键词
Hey there, let's tackle this substring keyword highlighting issue in PDFClown. The core goal is to make sure longer, complete keywords (like just.ETS or Test.ETS) get highlighted instead of their shorter substrings (like ETS), right? Here's a step-by-step approach to make this work smoothly:
1. Preprocess Keywords to Prioritize Longer Terms
First, we need to sort our keyword list by length in descending order. This ensures that longer, complete terms are checked first, so they don't get overshadowed by their shorter substrings.
import java.util.Arrays; import java.util.List; List<String> keywords = Arrays.asList("ETS", "just.ETS", "Test.ETS"); // Sort keywords from longest to shortest to prioritize full terms keywords.sort((a, b) -> Integer.compare(b.length(), a.length()));
2. Modify Highlighting Logic to Avoid Substring Overlaps
Next, we'll adjust the PDF text extraction and annotation logic in PDFClown. The key here is to track already highlighted regions so we don't re-highlight substrings within a fully matched keyword. We'll work with both the source ActualPdf and the output result PDF.
import org.pdfclown.documents.Page; import org.pdfclown.documents.interaction.annotations.HighlightAnnotation; import org.pdfclown.files.PdfDocument; import org.pdfclown.tools.TextExtractor; import org.pdfclown.util.math.Rectangle2D; import java.io.FileInputStream; import java.io.FileOutputStream; import java.util.HashSet; import java.util.Map; import java.util.Set; // Load source ActualPdf PdfDocument actualPdf = new PdfDocument(new FileInputStream("path/to/your/actual.pdf")); // Initialize result PDF for highlighted output PdfDocument resultPdf = new PdfDocument(); // Iterate through each page in the source PDF for (Page sourcePage : actualPdf.getPages()) { // Clone source page to result PDF Page resultPage = resultPdf.getPages().add(sourcePage.clone()); TextExtractor textExtractor = new TextExtractor(); Map<Rectangle2D, List<TextExtractor.TextChunk>> textChunks = textExtractor.extract(sourcePage); // Track regions already highlighted to prevent substring overlaps Set<Rectangle2D> highlightedRegions = new HashSet<>(); for (String keyword : keywords) { for (Map.Entry<Rectangle2D, List<TextExtractor.TextChunk>> entry : textChunks.entrySet()) { Rectangle2D chunkBounds = entry.getKey(); List<TextExtractor.TextChunk> chunks = entry.getValue(); String chunkText = TextExtractor.TextChunk.toString(chunks); if (chunkText.contains(keyword)) { // Calculate exact bounds of the keyword within the text chunk Rectangle2D keywordBounds = calculateKeywordBounds(chunkBounds, chunks, keyword); // Skip if this region is already covered by a longer keyword boolean isOverlapped = highlightedRegions.stream() .anyMatch(region -> region.intersects(keywordBounds)); if (!isOverlapped) { // Add yellow highlight annotation to result page HighlightAnnotation highlight = new HighlightAnnotation(resultPage, keywordBounds); highlight.setColor(new org.pdfclown.documents.contents.colorSpaces.DeviceRGBColor(1f, 1f, 0f)); resultPage.getAnnotations().add(highlight); // Add popup with associated measurement value addMeasurementPopup(resultPage, keywordBounds, getMeasurementValue(keyword)); // Mark this region as highlighted highlightedRegions.add(keywordBounds); } } } } } // Save the final highlighted PDF resultPdf.save(new FileOutputStream("path/to/your/highlighted_result.pdf"), org.pdfclown.files.SerializationModeEnum.STANDARD);
Helper Methods (You'll Need These)
calculateKeywordBounds: Computes the exact rectangular area of the keyword within the text chunk, using character positions and widths.addMeasurementPopup: Creates a popup annotation attached to the highlight, displaying the associated measurement value for the keyword.getMeasurementValue: Retrieves the measurement from a pre-defined mapping table (see step 4).
3. Configure PDFClown Jar Dependencies
Make sure your project has the correct PDFClown libraries set up:
- If using Maven, add this to your
pom.xml:
<dependency> <groupId>org.pdfclown</groupId> <artifactId>pdfclown-lib</artifactId> <version>0.1.2.3</version> </dependency>
- For manual setup, download the
pdfclown-lib.jarand add it to your project'slibdirectory, then configure your IDE to include it as a library dependency.
4. Use a Mapping Table for Keyword-Measurement Pairs
To link full keywords to their corresponding popup measurement values, maintain a hash map. This ensures we only associate measurements with valid, full keywords (not substrings):
import java.util.HashMap; import java.util.Map; Map<String, String> keywordMeasurementMap = new HashMap<>(); keywordMeasurementMap.put("just.ETS", "15.2 cm"); keywordMeasurementMap.put("Test.ETS", "9.7 mm"); // Note: We don't include "ETS" here since it's a substring and shouldn't be highlighted // Inside getMeasurementValue method: private String getMeasurementValue(String keyword) { return keywordMeasurementMap.getOrDefault(keyword, "No measurement available"); }
内容的提问来源于stack exchange,提问作者Seshadri

