You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:使用JSoup解析HTML表格数据

How to Extract Table Column Data with JSoup

Let’s break down exactly how to pull table column data using JSoup, plus tackle the large file and mixed style concerns you mentioned.

First: Efficiently Parse Your 5MB HTML File

Since your file is quite large, it’s smart to avoid unnecessary memory bloat. JSoup handles this well out of the box, but here’s the most reliable way to load it:

File input = new File("your-large-document.html");
Document doc = Jsoup.parse(input, "UTF-8"); // Uses file system directly, no extra string copies

If you hit memory errors, run your JVM with increased heap space (e.g., -Xmx2G) to give it more room to work with.

Step 1: Target the Right Table

First, identify your table using a precise selector. Use an ID if available (it’s the fastest and most reliable):

Element targetTable = doc.select("table#product-data-table").first(); // Replace with your table's ID

If no ID exists, narrow it down using parent elements or class names:

Element targetTable = doc.select("div.content-wrapper table.styled-data-table").first();

Pro tip: Use your browser’s dev tools (right-click the table → Inspect → Copy CSS selector) to get a perfect selector every time.

Step 2: Extract Rows and Columns

Once you have the table, loop through rows and pull column data. Here’s a complete, reusable example:

import org.jsoup.Jsoup;
import org.jsoup.nodes.Document;
import org.jsoup.nodes.Element;
import org.jsoup.select.Elements;
import java.io.File;
import java.io.IOException;

public class TableDataExtractor {
    public static void main(String[] args) throws IOException {
        // Load the large HTML file
        File input = new File("your-file.html");
        Document doc = Jsoup.parse(input, "UTF-8");

        // Grab the target table
        Element table = doc.select("table[data-table-type='inventory']").first();
        if (table == null) {
            System.err.println("Couldn't find the target table!");
            return;
        }

        // Iterate through each row (skip header if needed by using tbody tr)
        Elements rows = table.select("tr");
        for (int i = 1; i < rows.size(); i++) { // Start at 1 to skip header row
            Element row = rows.get(i);
            Elements columns = row.select("td");

            // Extract data from specific columns (adjust indices to match your table)
            String itemId = columns.get(0).text().trim();
            String itemName = columns.get(1).select("span.product-name").text().trim(); // Target nested element
            String stockCount = columns.get(2).text().trim();

            // Do something with the data (print, store in a list, etc.)
            System.out.printf("Item ID: %s | Name: %s | Stock: %s%n", itemId, itemName, stockCount);
        }
    }
}

Handling Mixed CSS/Inline Styles

If some columns are hidden or styled to obscure content, add checks to filter those out:

for (Element cell : columns) {
    // Skip cells hidden via CSS class or inline style
    if (cell.hasClass("hidden-column") || cell.attr("style").contains("display:none")) {
        continue;
    }
    // Process visible cell content here
}

Pro Tips for Large Files

  • Filter Early: Instead of parsing the entire document, use selectors to grab only the section containing your table first:
    Element tableContainer = doc.select("div.table-container").first();
    Element table = tableContainer.select("table").first();
    
  • Stream Parsing (Advanced): For ultra-large files, use JSoup’s XmlTreeBuilder with a BufferedReader to parse incrementally, though this is only necessary if standard parsing hits memory limits.
  • Avoid Unnecessary Data: Skip parsing elements you don’t need (like scripts or styles) by using doc.select("table") directly instead of loading the full document into memory.

Troubleshooting Common Issues

  • Empty Cell Text: If a cell’s content is in a nested element (like <span> or <div>), target that element specifically instead of using text() on the <td> itself.
  • Inconsistent Rows: If some rows have fewer columns, add a check before accessing indices:
    if (columns.size() >= 3) {
        String stockCount = columns.get(2).text().trim();
    }
    

内容的提问来源于stack exchange,提问作者Ziggy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 03:40:14