如何使用Apache POI XWPF获取docx文档绘图中的形状
从Apache POI的CTDrawing对象中提取Word形状的实现方案
问题场景
持有包含文本与若干形状的docx文档,已通过Apache POI成功获取文档内的CTDrawing绘图对象,可识别到inline(内嵌)、anchor(锚定)两类绘图节点下存在形状数据,但无法正确解析提取Shape对象,现有解析代码如下:
try (FileInputStream fis = new FileInputStream(file)) { document = new XWPFDocument(fis); parseDocument(document); }
核心解析逻辑代码:
public void parseDocument(XWPFDocument document) throws Exception { Iterator<XWPFParagraph> it = document.getParagraphs().iterator(); while (it.hasNext()) { XWPFParagraph paragraph = it.next(); List<XWPFRun> runs = paragraph.getRuns(); Iterator<XWPFRun> it2 = runs.iterator(); while (it2.hasNext()) { XWPFRun run = it2.next(); List<CTDrawing> drawings = getAllDrawings(run); if (!drawings.isEmpty()) { System.out.println("Found " + drawings.size() + " drawings"); Iterator<CTDrawing> it3 = drawings.iterator(); while (it3.hasNext()) { CTDrawing drawing = it3.next(); visitDrawing(drawing); } } } } } private void visitDrawing(CTDrawing drawing) { List<CTInline> inlines = drawing.getInlineList(); Iterator<CTInline> it = inlines.iterator(); while (it.hasNext()) { CTInline inline = it.next(); CTGraphicalObject graphic = inline.getGraphic(); if (graphic != null) { CTGraphicalObjectData graphicData = graphic.getGraphicData(); if (graphicData != null) { XmlCursor c = graphicData.newCursor(); c.selectPath("./*"); while (c.toNextSelection()) { XmlObject o = c.getObject(); System.out.println(o); } } } } List<CTAnchor> anchors = drawing.getAnchorList(); Iterator<CTAnchor> it2 = anchors.iterator(); while (it2.hasNext()) { CTAnchor anchor = it2.next(); CTGraphicalObject graphic = anchor.getGraphic(); if (graphic != null) { CTGraphicalObjectData graphicData = graphic.getGraphicData(); if (graphicData != null) { XmlCursor c = graphicData.newCursor(); c.selectPath("./*"); while (c.toNextSelection()) { XmlObject o = c.getObject(); System.out.println(o); } } } } } private static List<CTDrawing> getAllDrawings(XWPFRun run) throws Exception { CTR ctR = run.getCTR(); XmlCursor cursor = ctR.newCursor(); cursor.selectPath("declare namespace w='http://schemas.openxmlformats.org/wordprocessingml/2006/main' .//*/w:drawing"); List<CTDrawing> drawings = new ArrayList<>(); while (cursor.hasNextSelection()) { cursor.toNextSelection(); XmlObject obj = cursor.getObject(); CTDrawing drawing = CTDrawing.Factory.parse(obj.newInputStream()); drawings.add(drawing); } return drawings; }
实现步骤
1. 替换依赖
首先确保项目引入的是完整版POI OOXML依赖,精简版poi-ooxml-lite裁剪了形状相关的XML绑定类,会出现类找不到的问题,版本和当前使用的POI保持一致即可,Maven配置示例:
<dependency> <groupId>org.apache.poi</groupId> <artifactId>poi-ooxml-full</artifactId> <version>5.2.5</version> </dependency>
2. 修改绘图节点解析逻辑
Word中插入的普通自选图形、文本框、艺术字都存储在graphicData节点下,对应http://schemas.microsoft.com/office/word/2010/wordprocessingShape命名空间的CTShape节点,遍历子节点时直接做类型判断即可拿到形状对象,替换原有visitDrawing方法及新增解析逻辑如下:
// 统一处理内嵌、锚定绘图的图形数据 private void visitDrawing(CTDrawing drawing) { for (CTInline inline : drawing.getInlineList()) { extractShapesFromGraphic(inline.getGraphic()); } for (CTAnchor anchor : drawing.getAnchorList()) { extractShapesFromGraphic(anchor.getGraphic()); } } private void extractShapesFromGraphic(CTGraphicalObject graphic) { if (graphic == null) return; CTGraphicalObjectData graphicData = graphic.getGraphicData(); if (graphicData == null) return; try (XmlCursor cursor = graphicData.newCursor()) { cursor.selectPath("./*"); while (cursor.toNextSelection()) { XmlObject node = cursor.getObject(); // 优先使用POI封装的高层Shape API,使用更简单 if (node instanceof XWPFShape) { XWPFShape shape = (XWPFShape) node; // 读取基础属性 System.out.printf("找到形状,名称:%s,类型:%s%n", shape.getShapeName(), shape.getShapeType()); // 读取形状内的文本内容 if (shape instanceof XWPFShapeBase) { List<XWPFParagraph> paragraphs = ((XWPFShapeBase) shape).getTextParagraphs(); for (XWPFParagraph para : paragraphs) { System.out.println("形状内文本:" + para.getText()); } } } // 兼容低版本POI无高层API的场景,直接操作底层XML对象 else if (node instanceof com.microsoft.schemas.office.word.x2010.wordprocessingShape.CTShape) { var ctShape = (com.microsoft.schemas.office.word.x2010.wordprocessingShape.CTShape) node; // 读取基础属性 var shapePr = ctShape.getCnVPr(); System.out.printf("找到底层形状对象,ID:%d,名称:%s%n", shapePr.getId(), shapePr.getName()); // 读取形状内文本 CTTextBody textBody = ctShape.getTxBody(); if (textBody != null) { for (CTParagraph para : textBody.getPList()) { StringBuilder text = new StringBuilder(); for (CTRun run : para.getRList()) { text.append(run.getT()); } System.out.println("形状内文本:" + text); } } // 还可以读取位置、大小、填充色、边框等所有属性,直接调用ctShape对应的get方法即可 } } } }
注意事项
- 如果需要解析组合形状,拿到CTShape对象后递归遍历其下的
grpSp子节点即可,逻辑和上述解析流程一致。 - 高层XWPFShape API封装了常用属性读取方法,开发效率更高;如果需要读取特殊属性(比如形状的自定义XML属性、特效参数),直接操作底层CTShape对象即可拿到全量数据。
- 原有
getAllDrawings方法不需要修改,已经可以正确拿到所有绘图节点。
内容的提问来源于stack exchange,提问作者Hervé Girod
相关产品推荐
相关产品推荐

