You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从HWPFDocument提取2003格式Word文档的超链接?

从HWPFDocument提取超链接的解决方案

由于HWPF(处理旧版.doc格式Word文档)没有像XWPF那样提供直接的getHyperlinks() API,需要通过解析文档的字段或字符运行(CharacterRun)来提取超链接,以下是两种可行方案:

方案一:遍历文档字段提取

超链接在HWPF中以FIELD_HYPERLINK类型的字段存在,可通过遍历所有字段筛选并提取:

WordExtractor wordExtractor = null;
          
if(inFilePath == null || inFilePath.equals("")) {
    throw new Exception("Invalid file path");
}
HWPFDocument document = null;
InputStream is;
File file = new File(inFilePath);
is = new FileInputStream(file);

document = new HWPFDocument(is);

// 开始提取超链接
FieldIterator fieldIterator = document.getRange().getFields();
while (fieldIterator.hasNext()) {
    Field field = fieldIterator.next();
    if (field.getType() == FieldType.FIELD_HYPERLINK) {
        // 提取目标URL
        String fieldCode = field.getFieldCode();
        String url = fieldCode.substring(fieldCode.indexOf("HYPERLINK") + 9).trim();
        // 去除URL两端的引号
        if (url.startsWith("\"") && url.endsWith("\"")) {
            url = url.substring(1, url.length() - 1);
        }
        // 提取显示文本
        String displayText = field.getResult();
        // 输出或处理超链接
        System.out.printf("超链接文本:%s,目标URL:%s%n", displayText, url);
    }
}

wordExtractor = new WordExtractor(document);
// 后续代码...

方案二:遍历段落的CharacterRun提取

每个超链接会关联到对应的CharacterRun,可通过遍历段落中的字符运行获取:

WordExtractor wordExtractor = null;
          
if(inFilePath == null || inFilePath.equals("")) {
    throw new Exception("Invalid file path");
}
HWPFDocument document = null;
InputStream is;
File file = new File(inFilePath);
is = new FileInputStream(file);

document = new HWPFDocument(is);

// 开始提取超链接
Range documentRange = document.getRange();
for (int paraIndex = 0; paraIndex < documentRange.numParagraphs(); paraIndex++) {
    Paragraph para = documentRange.getParagraph(paraIndex);
    for (int runIndex = 0; runIndex < para.numCharacterRuns(); runIndex++) {
        CharacterRun run = para.getCharacterRun(runIndex);
        Hyperlink hyperlink = run.getHyperlink();
        if (hyperlink != null) {
            String url = hyperlink.getAddress();
            String displayText = run.text();
            // 输出或处理超链接
            System.out.printf("超链接文本:%s,目标URL:%s%n", displayText, url);
        }
    }
}

wordExtractor = new WordExtractor(document);
// 后续代码...

注意事项

  • 方案一基于字段解析,能更准确获取超链接的完整显示文本(即使文本跨多个CharacterRun);
  • 方案二更直观,但如果超链接的显示文本被拆分到多个CharacterRun中,可能需要额外处理拼接。

内容的提问来源于stack exchange,提问作者Code Trickle

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 09:22:43