如何使用Apache POI解析XML中HTML表格并生成Word文档?
解决方案
1. 还原转义的HTML内容
你通过element.item(n).getChildNodes().item(0).getNodeValue()获取到的是HTML转义字符串(比如<table>这类转义后的标签),第一步需要把它还原成真实的HTML标签。
Java中可以用commons-text工具类快速处理:
// 获取转义后的原始字符串 String escapedHtml = element.item(n).getChildNodes().item(0).getNodeValue().trim(); // 还原为可解析的HTML String unescapedHtml = org.apache.commons.text.StringEscapeUtils.unescapeHtml4(escapedHtml);
2. 解析HTML提取表格结构
使用Jsoup库解析还原后的HTML,方便提取表格的行、列及内容:
依赖引入(Maven)
<dependency> <groupId>org.jsoup</groupId> <artifactId>jsoup</artifactId> <version>1.17.2</version> </dependency>
解析代码示例
// 解析HTML文档 org.jsoup.nodes.Document htmlDoc = Jsoup.parse(unescapedHtml); // 定位目标表格(根据class筛选) org.jsoup.nodes.Element table = htmlDoc.selectFirst("table.conTable"); if (table != null) { // 遍历表格行 org.jsoup.select.Elements rows = table.select("tbody tr"); for (org.jsoup.nodes.Element row : rows) { // 遍历每行的单元格 org.jsoup.select.Elements cells = row.select("td.confluenceTd"); for (org.jsoup.nodes.Element cell : cells) { // 提取单元格内的列表项 org.jsoup.select.Elements productItems = cell.select("ul li"); for (org.jsoup.nodes.Element item : productItems) { String productName = item.text().trim(); // 此处可暂存数据,或直接写入Word } } } }
3. 用Apache POI生成对应Word表格
根据提取到的表格结构,使用POI的XWPF组件创建Word表格并填充内容:
// 创建Word文档对象 org.apache.poi.xwpf.usermodel.XWPFDocument doc = new org.apache.poi.xwpf.usermodel.XWPFDocument(); // 创建对应行列数的表格 org.apache.poi.xwpf.usermodel.XWPFTable wordTable = doc.createTable(rows.size(), cells.size()); int rowIndex = 0; for (org.jsoup.nodes.Element htmlRow : rows) { org.apache.poi.xwpf.usermodel.XWPFTableRow wordRow = wordTable.getRow(rowIndex); org.jsoup.select.Elements htmlCells = htmlRow.select("td.confluenceTd"); int cellIndex = 0; for (org.jsoup.nodes.Element htmlCell : htmlCells) { org.apache.poi.xwpf.usermodel.XWPFTableCell wordCell = wordRow.getCell(cellIndex); // 拼接单元格内的列表内容(如需保留项目符号,可通过POI创建段落列表) StringBuilder cellContent = new StringBuilder(); org.jsoup.select.Elements items = htmlCell.select("ul li"); for (org.jsoup.nodes.Element item : items) { cellContent.append("• ").append(item.text().trim()).append("\n"); } wordCell.setText(cellContent.toString().trim()); cellIndex++; } rowIndex++; } // 保存Word文档 try (java.io.FileOutputStream out = new java.io.FileOutputStream("product_table.docx")) { doc.write(out); }
注意事项
- 需确保引入完整依赖:除POI的
poi、poi-ooxml,还要添加commons-text和jsoup; - 若HTML存在语法错误(如你提供的XML里的
Product3;/li>),Jsoup会自动修复,但建议提前做文本清理; - 如需保留Word内的原生项目符号列表,可通过
XWPFParagraph和XWPFRun手动创建,而非直接拼接文本。
内容的提问来源于stack exchange,提问作者Jeet
相关产品推荐
相关产品推荐

