Java中使用Aspose Words获取Word内嵌文档数据的方法
如何从Aspose Words识别为Shape的内嵌文档中提取内容
使用Aspose Words直接处理
Aspose Words中,这类内嵌文档本质是OLE对象,被封装在Shape节点内。你可以通过以下步骤提取内容:
- 先验证Shape是否包含目标OLE对象:检查
Shape.OleFormat属性非空,且OleFormat.ProgId匹配Word文档类型(如Word.Document.12)。 - 提取内嵌文档数据:通过
OleFormat.GetOleEntry("Contents")获取字节流,再将字节流加载为新的Aspose.Words.Document对象,即可正常读取其中内容。
.NET示例代码
foreach (Shape shape in doc.GetChildNodes(NodeType.Shape, true)) { if (shape.OleFormat != null && shape.OleFormat.ProgId.StartsWith("Word.Document")) { // 获取内嵌文档的字节流 byte[] oleBytes = shape.OleFormat.GetOleEntry("Contents"); // 加载为独立的Document对象 using (MemoryStream ms = new MemoryStream(oleBytes)) { Document embeddedDoc = new Document(ms); // 提取内嵌文档文本(可按需处理其他内容) string embeddedText = embeddedDoc.GetText(); Console.WriteLine(embeddedText); } } }
Java示例代码
for (Shape shape : (Iterable<Shape>) doc.getChildNodes(NodeType.SHAPE, true)) { if (shape.getOleFormat() != null && shape.getOleFormat().getProgId().startsWith("Word.Document")) { byte[] oleBytes = shape.getOleFormat().getOleEntry("Contents"); ByteArrayInputStream bais = new ByteArrayInputStream(oleBytes); Document embeddedDoc = new Document(bais); // 提取内嵌文档文本 String embeddedText = embeddedDoc.getText(); System.out.println(embeddedText); } }
其他可选库
如果不想使用Aspose Words,也可以选择这些工具:
- Apache POI(Java生态):通过
XWPFDocument配合POIOLE2TextExtractor解析OLE包,进而提取内嵌Word文档内容,但对复杂格式的兼容性略逊于Aspose。 - DocX(.NET生态):轻量级Word处理库,适合简单场景的内嵌文档提取,但对复杂OLE对象的支持不够完善。
注意:.doc与.docx格式的OLE对象存储结构略有差异,测试时需覆盖两种格式的文件。
内容的提问来源于stack exchange,提问作者lemon chow
相关产品推荐
相关产品推荐

