You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java中使用Aspose Words获取Word内嵌文档数据的方法

如何从Aspose Words识别为Shape的内嵌文档中提取内容

使用Aspose Words直接处理

Aspose Words中,这类内嵌文档本质是OLE对象,被封装在Shape节点内。你可以通过以下步骤提取内容:

  • 先验证Shape是否包含目标OLE对象:检查Shape.OleFormat属性非空,且OleFormat.ProgId匹配Word文档类型(如Word.Document.12)。
  • 提取内嵌文档数据:通过OleFormat.GetOleEntry("Contents")获取字节流,再将字节流加载为新的Aspose.Words.Document对象,即可正常读取其中内容。

.NET示例代码

foreach (Shape shape in doc.GetChildNodes(NodeType.Shape, true))
{
    if (shape.OleFormat != null && shape.OleFormat.ProgId.StartsWith("Word.Document"))
    {
        // 获取内嵌文档的字节流
        byte[] oleBytes = shape.OleFormat.GetOleEntry("Contents");
        // 加载为独立的Document对象
        using (MemoryStream ms = new MemoryStream(oleBytes))
        {
            Document embeddedDoc = new Document(ms);
            // 提取内嵌文档文本(可按需处理其他内容)
            string embeddedText = embeddedDoc.GetText();
            Console.WriteLine(embeddedText);
        }
    }
}

Java示例代码

for (Shape shape : (Iterable<Shape>) doc.getChildNodes(NodeType.SHAPE, true)) {
    if (shape.getOleFormat() != null && shape.getOleFormat().getProgId().startsWith("Word.Document")) {
        byte[] oleBytes = shape.getOleFormat().getOleEntry("Contents");
        ByteArrayInputStream bais = new ByteArrayInputStream(oleBytes);
        Document embeddedDoc = new Document(bais);
        // 提取内嵌文档文本
        String embeddedText = embeddedDoc.getText();
        System.out.println(embeddedText);
    }
}

其他可选库

如果不想使用Aspose Words,也可以选择这些工具:

  • Apache POI(Java生态):通过XWPFDocument配合POIOLE2TextExtractor解析OLE包,进而提取内嵌Word文档内容,但对复杂格式的兼容性略逊于Aspose。
  • DocX(.NET生态):轻量级Word处理库,适合简单场景的内嵌文档提取,但对复杂OLE对象的支持不够完善。

注意:.doc与.docx格式的OLE对象存储结构略有差异,测试时需覆盖两种格式的文件。

内容的提问来源于stack exchange,提问作者lemon chow

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 10:57:15