You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Jsoup无法解析JS动态生成的线索表格,寻求解决方案

Solution for Parsing Dynamically Loaded Tables with Jsoup & Selenium

Got it, let's tackle this problem head-on. The core issue here is that Jsoup is a static HTML parser—it can't execute JavaScript or wait for AJAX-loaded content to render. That's why your target table (populated by the UpdateDiv() AJAX function) is invisible to Jsoup when you try to parse the raw page source directly.

Selenium is the perfect tool to bridge this gap. You don't need to worry about Selenium "affecting" Jsoup's parsing—instead, Selenium will fully render the page (including executing the UpdateDiv() function and loading the table via AJAX), then you can pass the fully rendered HTML to Jsoup, or even extract the data directly using Selenium. Here are two solid approaches:


Approach 1: Use Selenium to Get Rendered HTML, Then Parse with Jsoup

This method leverages Selenium to handle the dynamic rendering, then hands off the complete HTML to Jsoup for the parsing you're already comfortable with.

Step-by-Step Implementation (Java Example)

First, make sure you have the necessary dependencies (Selenium and Jsoup) in your project. For Maven, add these to your pom.xml:

<dependencies>
    <!-- Selenium -->
    <dependency>
        <groupId>org.seleniumhq.selenium</groupId>
        <artifactId>selenium-java</artifactId>
        <version>4.15.0</version>
    </dependency>
    <!-- Jsoup -->
    <dependency>
        <groupId>org.jsoup</groupId>
        <artifactId>jsoup</artifactId>
        <version>1.17.2</version>
    </dependency>
</dependencies>

Then, use this code to render the page and parse the table:

import org.jsoup.Jsoup;
import org.jsoup.nodes.Document;
import org.openqa.selenium.WebDriver;
import org.openqa.selenium.chrome.ChromeDriver;
import org.openqa.selenium.chrome.ChromeOptions;
import org.openqa.selenium.support.ui.WebDriverWait;
import org.openqa.selenium.support.ui.ExpectedConditions;
import java.time.Duration;

public class DynamicLeadParser {
    public static void main(String[] args) {
        // Configure ChromeDriver (ensure it matches your Chrome browser version)
        System.setProperty("webdriver.chrome.driver", "path/to/your/chromedriver");

        // Use headless mode to avoid opening a visible browser window (optional but great for servers)
        ChromeOptions options = new ChromeOptions();
        options.addArguments("--headless=new");

        WebDriver driver = new ChromeDriver(options);
        WebDriverWait wait = new WebDriverWait(driver, Duration.ofSeconds(15));

        try {
            // Load your target page
            driver.get("your-target-page-url-here");

            // Execute the UpdateDiv() function to trigger the AJAX table load
            driver.executeScript("UpdateDiv();");

            // Wait for the target table to appear (more reliable than hardcoding sleep)
            wait.until(ExpectedConditions.presenceOfElementLocated(org.openqa.selenium.By.cssSelector("table.none")));

            // Grab the fully rendered page source
            String renderedHtml = driver.getPageSource();

            // Now parse this HTML with Jsoup
            Document doc = Jsoup.parse(renderedHtml);

            // Extract your lead data from the table
            var leadTable = doc.selectFirst("table.none");
            String fullName = leadTable.select("tr.tm_tt_body td.typedata1[colspan=3]").first().text().trim();
            String phoneNumber = leadTable.select("b:contains(222-222-2222)").text().trim();
            String email = leadTable.select("#ld_email").attr("href").replace("mailto:", "").split("\\?subject")[0].trim();

            // Print or process the extracted data
            System.out.println("Lead Name: " + fullName);
            System.out.println("Lead Phone: " + phoneNumber);
            System.out.println("Lead Email: " + email);

        } catch (Exception e) {
            e.printStackTrace();
        } finally {
            // Always quit the driver to clean up resources
            driver.quit();
        }
    }
}

Approach 2: Extract Data Directly with Selenium

Since you're already using Selenium to render the page, you can skip Jsoup entirely and extract the data directly using Selenium's element locators. This is simpler if you don't need Jsoup's specific parsing features.

Example Code Snippet

import org.openqa.selenium.WebDriver;
import org.openqa.selenium.chrome.ChromeDriver;
import org.openqa.selenium.chrome.ChromeOptions;
import org.openqa.selenium.WebElement;
import org.openqa.selenium.support.ui.WebDriverWait;
import org.openqa.selenium.support.ui.ExpectedConditions;
import java.time.Duration;

public class SeleniumDirectExtractor {
    public static void main(String[] args) {
        System.setProperty("webdriver.chrome.driver", "path/to/your/chromedriver");
        ChromeOptions options = new ChromeOptions();
        options.addArguments("--headless=new");

        WebDriver driver = new ChromeDriver(options);
        WebDriverWait wait = new WebDriverWait(driver, Duration.ofSeconds(15));

        try {
            driver.get("your-target-page-url-here");
            driver.executeScript("UpdateDiv();");

            // Wait for the table to load
            WebElement leadTable = wait.until(ExpectedConditions.presenceOfElementLocated(org.openqa.selenium.By.cssSelector("table.none")));

            // Extract data directly via Selenium locators
            WebElement nameElement = leadTable.findElement(org.openqa.selenium.By.cssSelector("tr.tm_tt_body td.typedata1[colspan=3]"));
            String fullName = nameElement.getText().trim();

            WebElement phoneElement = leadTable.findElement(org.openqa.selenium.By.xpath("//a[contains(text(), '222-222-2222')]"));
            String phoneNumber = phoneElement.getText().trim();

            WebElement emailElement = leadTable.findElement(org.openqa.selenium.By.id("ld_email"));
            String email = emailElement.getAttribute("href").replace("mailto:", "").split("\\?subject")[0].trim();

            System.out.println("Lead Name: " + fullName);
            System.out.println("Lead Phone: " + phoneNumber);
            System.out.println("Lead Email: " + email);

        } catch (Exception e) {
            e.printStackTrace();
        } finally {
            driver.quit();
        }
    }
}

Key Notes

  • Browser Driver Compatibility: Make sure your ChromeDriver version matches your installed Chrome browser version (download from ChromeDriver official site—but wait, no external links allowed, so just remember to get the matching version).
  • Reliable Waiting: Use WebDriverWait with explicit conditions (like waiting for the table to exist) instead of Thread.sleep()—this makes your code more robust to varying network speeds.
  • Headless Mode: The --headless=new flag runs Chrome without a visible window, which is ideal for running scripts on servers or in background processes.

内容的提问来源于stack exchange,提问作者Brandon Woodruff

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:18:26