如何从HWPFDocument提取2003格式Word文档的超链接?
从HWPFDocument提取超链接的解决方案
由于HWPF(处理旧版.doc格式Word文档)没有像XWPF那样提供直接的getHyperlinks() API,需要通过解析文档的字段或字符运行(CharacterRun)来提取超链接,以下是两种可行方案:
方案一:遍历文档字段提取
超链接在HWPF中以FIELD_HYPERLINK类型的字段存在,可通过遍历所有字段筛选并提取:
WordExtractor wordExtractor = null; if(inFilePath == null || inFilePath.equals("")) { throw new Exception("Invalid file path"); } HWPFDocument document = null; InputStream is; File file = new File(inFilePath); is = new FileInputStream(file); document = new HWPFDocument(is); // 开始提取超链接 FieldIterator fieldIterator = document.getRange().getFields(); while (fieldIterator.hasNext()) { Field field = fieldIterator.next(); if (field.getType() == FieldType.FIELD_HYPERLINK) { // 提取目标URL String fieldCode = field.getFieldCode(); String url = fieldCode.substring(fieldCode.indexOf("HYPERLINK") + 9).trim(); // 去除URL两端的引号 if (url.startsWith("\"") && url.endsWith("\"")) { url = url.substring(1, url.length() - 1); } // 提取显示文本 String displayText = field.getResult(); // 输出或处理超链接 System.out.printf("超链接文本:%s,目标URL:%s%n", displayText, url); } } wordExtractor = new WordExtractor(document); // 后续代码...
方案二:遍历段落的CharacterRun提取
每个超链接会关联到对应的CharacterRun,可通过遍历段落中的字符运行获取:
WordExtractor wordExtractor = null; if(inFilePath == null || inFilePath.equals("")) { throw new Exception("Invalid file path"); } HWPFDocument document = null; InputStream is; File file = new File(inFilePath); is = new FileInputStream(file); document = new HWPFDocument(is); // 开始提取超链接 Range documentRange = document.getRange(); for (int paraIndex = 0; paraIndex < documentRange.numParagraphs(); paraIndex++) { Paragraph para = documentRange.getParagraph(paraIndex); for (int runIndex = 0; runIndex < para.numCharacterRuns(); runIndex++) { CharacterRun run = para.getCharacterRun(runIndex); Hyperlink hyperlink = run.getHyperlink(); if (hyperlink != null) { String url = hyperlink.getAddress(); String displayText = run.text(); // 输出或处理超链接 System.out.printf("超链接文本:%s,目标URL:%s%n", displayText, url); } } } wordExtractor = new WordExtractor(document); // 后续代码...
注意事项
- 方案一基于字段解析,能更准确获取超链接的完整显示文本(即使文本跨多个CharacterRun);
- 方案二更直观,但如果超链接的显示文本被拆分到多个CharacterRun中,可能需要额外处理拼接。
内容的提问来源于stack exchange,提问作者Code Trickle
相关产品推荐
相关产品推荐

