使用Java 8解析含rowspan和colspan的复杂HTML表格为二维数组
处理带rowspan和colspan的HTML表格解析问题
问题描述
我需要解析一个包含rowspan和colspan属性的HTML表格,参考了Stack Overflow上的方案,但其中这段代码会抛出IndexOutOfBoundsException:
row.add(idx, rowspan - 1 == 0 ? (Element) td.removeAttr("rowspan") : td.attr("rowspan", String.valueOf(rowspan - 1)));
求能同时处理rowspan和colspan,将表格解析为二维数组的解决思路。
示例表格代码:
<table border="1"> <tr> <td rowspan="3" colspan="2">head1</td> <td rowspan="2" colspan="2">head2</td> <td colspan="6">head3</td> </tr> <tr> <td colspan="2">head3-1</td> <td colspan="2">head3-2</td> <td colspan="2">head3-3</td> </tr> <tr> <td>sub_header1</td> <td>sub_header2</td> <td>sub_header3</td> <td>sub_header4</td> <td>sub_header5</td> <td>sub_header6</td> <td>sub_header7</td> <td>sub_header8</td> </tr> <tr> <td>content1</td> <td>content2</td> <td>content3</td> <td>content4</td> <td>content5</td> <td>content6</td> <td>content7</td> <td>content8</td> <td>content9</td> <td>content10</td> </tr> </table>
核心解决思路
1. 维护跨行列的缓存列表
不要直接在当前行操作下一行元素,而是用缓存记录需要向下填充的元素:
- 遍历每行时,先从缓存中取出需要填充到当前行的元素,填充对应位置后更新缓存(剩余行数减1,行数为0则移除)。
- 处理当前行的原始
<td>时,先跳过已被缓存填充的位置,再根据colspan填充当前行的连续列,rowspan>1时将元素加入后续行的缓存,标记填充范围和剩余行数。
2. 提前计算表格最大列数
因为colspan会让行的实际列数变化,先遍历所有行,计算展开所有colspan后的最大列数,作为二维数组的固定列长度,从根源避免索引越界。
3. 替换原代码的索引操作逻辑
原代码用row.add(idx, ...)会因idx超过当前行元素数量报错,正确做法是:
- 初始化当前行为固定长度的数组(用最大列数),或先填充占位符到指定长度,再用
set方法替换对应位置的元素,而非直接add。
代码实现示例(Java + Jsoup)
import org.jsoup.Jsoup; import org.jsoup.nodes.Document; import org.jsoup.nodes.Element; import org.jsoup.select.Elements; import java.util.ArrayList; import java.util.List; public class TableParser { public static List<List<String>> parseTable(String html) { Document doc = Jsoup.parse(html); Element table = doc.selectFirst("table"); Elements rows = table.select("tr"); // 计算表格最大列数(展开colspan后的总列数) int maxCols = 0; for (Element row : rows) { int colCount = 0; for (Element td : row.select("td, th")) { int colspan = td.hasAttr("colspan") ? Integer.parseInt(td.attr("colspan")) : 1; colCount += colspan; } maxCols = Math.max(maxCols, colCount); } // 缓存:每行待填充的元素,格式为[起始列, 结束列, 内容, 剩余行数] List<List<Object[]>> rowCache = new ArrayList<>(); List<List<String>> result = new ArrayList<>(); for (int rowIdx = 0; rowIdx < rows.size(); rowIdx++) { Element currentRow = rows.get(rowIdx); List<String> currentResultRow = new ArrayList<>(); // 初始化当前行为空占位符 for (int i = 0; i < maxCols; i++) { currentResultRow.add(""); } // 处理当前行的缓存元素 if (rowIdx < rowCache.size()) { List<Object[]> cacheItems = rowCache.get(rowIdx); for (Object[] item : cacheItems) { int startCol = (int) item[0]; int endCol = (int) item[1]; String content = (String) item[2]; // 填充对应列 for (int c = startCol; c <= endCol; c++) { currentResultRow.set(c, content); } // 更新剩余行数,若还有剩余则放入下一行缓存 int remainingRows = (int) item[3] - 1; if (remainingRows > 0) { if (rowIdx + 1 >= rowCache.size()) { rowCache.add(new ArrayList<>()); } rowCache.get(rowIdx + 1).add(new Object[]{startCol, endCol, content, remainingRows}); } } } // 处理当前行的原始td元素 int currentCol = 0; Elements tds = currentRow.select("td, th"); for (Element td : tds) { // 跳过已被缓存填充的列 while (currentCol < maxCols && !currentResultRow.get(currentCol).isEmpty()) { currentCol++; } if (currentCol >= maxCols) break; String content = td.text().trim(); int colspan = td.hasAttr("colspan") ? Integer.parseInt(td.attr("colspan")) : 1; int rowspan = td.hasAttr("rowspan") ? Integer.parseInt(td.attr("rowspan")) : 1; // 填充当前行的对应列 int endCol = currentCol + colspan - 1; for (int c = currentCol; c <= endCol && c < maxCols; c++) { currentResultRow.set(c, content); } // 处理rowspan,将元素加入后续行缓存 if (rowspan > 1) { int remainingRows = rowspan - 1; for (int r = rowIdx + 1; r <= rowIdx + remainingRows; r++) { if (r >= rowCache.size()) { rowCache.add(new ArrayList<>()); } rowCache.get(r).add(new Object[]{currentCol, endCol, content, remainingRows - (r - rowIdx)}); } } currentCol += colspan; } result.add(currentResultRow); } return result; } public static void main(String[] args) { String html = "<table border=\"1\">\n" + " <tr>\n" + " <td rowspan=\"3\" colspan=\"2\">head1</td>\n" + " <td rowspan=\"2\" colspan=\"2\">head2</td>\n" + " <td colspan=\"6\">head3</td>\n" + " </tr>\n" + " <tr>\n" + " <td colspan=\"2\">head3-1</td>\n" + " <td colspan=\"2\">head3-2</td>\n" + " <td colspan=\"2\">head3-3</td>\n" + " </tr>\n" + " <tr>\n" + " <td>sub_header1</td>\n" + " <td>sub_header2</td>\n" + " <td>sub_header3</td>\n" + " <td>sub_header4</td>\n" + " <td>sub_header5</td>\n" + " <td>sub_header6</td>\n" + " <td>sub_header7</td>\n" + " <td>sub_header8</td>\n" + " </tr>\n" + " <tr>\n" + " <td>content1</td> \n" + " <td>content2</td> \n" + " <td>content3</td> \n" + " <td>content4</td> \n" + " <td>content5</td> \n" + " <td>content6</td> \n" + " <td>content7</td> \n" + " <td>content8</td> \n" + " <td>content9</td> \n" + " <td>content10</td>\n" + " </tr>\n" + "</table>"; List<List<String>> tableData = parseTable(html); // 打印解析结果 for (List<String> row : tableData) { System.out.println(row); } } }
关键说明
- 先计算最大列数,彻底避免索引越界问题
- 缓存机制分离跨行列的填充逻辑,让每行的处理流程更清晰
- 处理
colspan时连续填充对应列,rowspan时将元素传递到后续行的缓存,逐行递减剩余填充行数
内容的提问来源于stack exchange,提问作者Noobgrammer
相关产品推荐
相关产品推荐

