You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Java 8解析含rowspan和colspan的复杂HTML表格为二维数组

处理带rowspan和colspan的HTML表格解析问题

问题描述

我需要解析一个包含rowspan和colspan属性的HTML表格,参考了Stack Overflow上的方案,但其中这段代码会抛出IndexOutOfBoundsException:

row.add(idx, rowspan - 1 == 0 ? (Element) td.removeAttr("rowspan") : td.attr("rowspan", String.valueOf(rowspan - 1)));

求能同时处理rowspan和colspan,将表格解析为二维数组的解决思路。

示例表格代码:

<table border="1">
  <tr>
    <td rowspan="3" colspan="2">head1</td>
    <td rowspan="2" colspan="2">head2</td>
    <td colspan="6">head3</td>
  </tr>
  <tr>
    <td colspan="2">head3-1</td>
    <td colspan="2">head3-2</td>
    <td colspan="2">head3-3</td>
  </tr>
  <tr>
    <td>sub_header1</td>
    <td>sub_header2</td>
    <td>sub_header3</td>
    <td>sub_header4</td>
    <td>sub_header5</td>
    <td>sub_header6</td>
    <td>sub_header7</td>
    <td>sub_header8</td>
  </tr>
  <tr>
    <td>content1</td>    
    <td>content2</td>    
    <td>content3</td>    
    <td>content4</td>    
    <td>content5</td>    
    <td>content6</td>    
    <td>content7</td>    
    <td>content8</td>    
    <td>content9</td>    
    <td>content10</td>
  </tr>
</table>

核心解决思路

1. 维护跨行列的缓存列表

不要直接在当前行操作下一行元素,而是用缓存记录需要向下填充的元素:

  • 遍历每行时,先从缓存中取出需要填充到当前行的元素,填充对应位置后更新缓存(剩余行数减1,行数为0则移除)。
  • 处理当前行的原始<td>时,先跳过已被缓存填充的位置,再根据colspan填充当前行的连续列,rowspan>1时将元素加入后续行的缓存,标记填充范围和剩余行数。

2. 提前计算表格最大列数

因为colspan会让行的实际列数变化,先遍历所有行,计算展开所有colspan后的最大列数,作为二维数组的固定列长度,从根源避免索引越界。

3. 替换原代码的索引操作逻辑

原代码用row.add(idx, ...)会因idx超过当前行元素数量报错,正确做法是:

  • 初始化当前行为固定长度的数组(用最大列数),或先填充占位符到指定长度,再用set方法替换对应位置的元素,而非直接add。

代码实现示例(Java + Jsoup)

import org.jsoup.Jsoup;
import org.jsoup.nodes.Document;
import org.jsoup.nodes.Element;
import org.jsoup.select.Elements;
import java.util.ArrayList;
import java.util.List;

public class TableParser {
    public static List<List<String>> parseTable(String html) {
        Document doc = Jsoup.parse(html);
        Element table = doc.selectFirst("table");
        Elements rows = table.select("tr");
        
        // 计算表格最大列数(展开colspan后的总列数)
        int maxCols = 0;
        for (Element row : rows) {
            int colCount = 0;
            for (Element td : row.select("td, th")) {
                int colspan = td.hasAttr("colspan") ? Integer.parseInt(td.attr("colspan")) : 1;
                colCount += colspan;
            }
            maxCols = Math.max(maxCols, colCount);
        }

        // 缓存:每行待填充的元素,格式为[起始列, 结束列, 内容, 剩余行数]
        List<List<Object[]>> rowCache = new ArrayList<>();
        List<List<String>> result = new ArrayList<>();

        for (int rowIdx = 0; rowIdx < rows.size(); rowIdx++) {
            Element currentRow = rows.get(rowIdx);
            List<String> currentResultRow = new ArrayList<>();
            // 初始化当前行为空占位符
            for (int i = 0; i < maxCols; i++) {
                currentResultRow.add("");
            }

            // 处理当前行的缓存元素
            if (rowIdx < rowCache.size()) {
                List<Object[]> cacheItems = rowCache.get(rowIdx);
                for (Object[] item : cacheItems) {
                    int startCol = (int) item[0];
                    int endCol = (int) item[1];
                    String content = (String) item[2];
                    // 填充对应列
                    for (int c = startCol; c <= endCol; c++) {
                        currentResultRow.set(c, content);
                    }
                    // 更新剩余行数,若还有剩余则放入下一行缓存
                    int remainingRows = (int) item[3] - 1;
                    if (remainingRows > 0) {
                        if (rowIdx + 1 >= rowCache.size()) {
                            rowCache.add(new ArrayList<>());
                        }
                        rowCache.get(rowIdx + 1).add(new Object[]{startCol, endCol, content, remainingRows});
                    }
                }
            }

            // 处理当前行的原始td元素
            int currentCol = 0;
            Elements tds = currentRow.select("td, th");
            for (Element td : tds) {
                // 跳过已被缓存填充的列
                while (currentCol < maxCols && !currentResultRow.get(currentCol).isEmpty()) {
                    currentCol++;
                }
                if (currentCol >= maxCols) break;

                String content = td.text().trim();
                int colspan = td.hasAttr("colspan") ? Integer.parseInt(td.attr("colspan")) : 1;
                int rowspan = td.hasAttr("rowspan") ? Integer.parseInt(td.attr("rowspan")) : 1;

                // 填充当前行的对应列
                int endCol = currentCol + colspan - 1;
                for (int c = currentCol; c <= endCol && c < maxCols; c++) {
                    currentResultRow.set(c, content);
                }

                // 处理rowspan,将元素加入后续行缓存
                if (rowspan > 1) {
                    int remainingRows = rowspan - 1;
                    for (int r = rowIdx + 1; r <= rowIdx + remainingRows; r++) {
                        if (r >= rowCache.size()) {
                            rowCache.add(new ArrayList<>());
                        }
                        rowCache.get(r).add(new Object[]{currentCol, endCol, content, remainingRows - (r - rowIdx)});
                    }
                }

                currentCol += colspan;
            }

            result.add(currentResultRow);
        }

        return result;
    }

    public static void main(String[] args) {
        String html = "<table border=\"1\">\n" +
                "  <tr>\n" +
                "    <td rowspan=\"3\" colspan=\"2\">head1</td>\n" +
                "    <td rowspan=\"2\" colspan=\"2\">head2</td>\n" +
                "    <td colspan=\"6\">head3</td>\n" +
                "  </tr>\n" +
                "  <tr>\n" +
                "    <td colspan=\"2\">head3-1</td>\n" +
                "    <td colspan=\"2\">head3-2</td>\n" +
                "    <td colspan=\"2\">head3-3</td>\n" +
                "  </tr>\n" +
                "  <tr>\n" +
                "    <td>sub_header1</td>\n" +
                "    <td>sub_header2</td>\n" +
                "    <td>sub_header3</td>\n" +
                "    <td>sub_header4</td>\n" +
                "    <td>sub_header5</td>\n" +
                "    <td>sub_header6</td>\n" +
                "    <td>sub_header7</td>\n" +
                "    <td>sub_header8</td>\n" +
                "  </tr>\n" +
                "  <tr>\n" +
                "    <td>content1</td>    \n" +
                "    <td>content2</td>    \n" +
                "    <td>content3</td>    \n" +
                "    <td>content4</td>    \n" +
                "    <td>content5</td>    \n" +
                "    <td>content6</td>    \n" +
                "    <td>content7</td>    \n" +
                "    <td>content8</td>    \n" +
                "    <td>content9</td>    \n" +
                "    <td>content10</td>\n" +
                "  </tr>\n" +
                "</table>";
        List<List<String>> tableData = parseTable(html);
        // 打印解析结果
        for (List<String> row : tableData) {
            System.out.println(row);
        }
    }
}

关键说明

  • 先计算最大列数,彻底避免索引越界问题
  • 缓存机制分离跨行列的填充逻辑,让每行的处理流程更清晰
  • 处理colspan时连续填充对应列,rowspan时将元素传递到后续行的缓存,逐行递减剩余填充行数

内容的提问来源于stack exchange,提问作者Noobgrammer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 19:50:28