Jsoup无法解析JS动态生成的线索表格,寻求解决方案
Got it, let's tackle this problem head-on. The core issue here is that Jsoup is a static HTML parser—it can't execute JavaScript or wait for AJAX-loaded content to render. That's why your target table (populated by the UpdateDiv() AJAX function) is invisible to Jsoup when you try to parse the raw page source directly.
Selenium is the perfect tool to bridge this gap. You don't need to worry about Selenium "affecting" Jsoup's parsing—instead, Selenium will fully render the page (including executing the UpdateDiv() function and loading the table via AJAX), then you can pass the fully rendered HTML to Jsoup, or even extract the data directly using Selenium. Here are two solid approaches:
Approach 1: Use Selenium to Get Rendered HTML, Then Parse with Jsoup
This method leverages Selenium to handle the dynamic rendering, then hands off the complete HTML to Jsoup for the parsing you're already comfortable with.
Step-by-Step Implementation (Java Example)
First, make sure you have the necessary dependencies (Selenium and Jsoup) in your project. For Maven, add these to your pom.xml:
<dependencies> <!-- Selenium --> <dependency> <groupId>org.seleniumhq.selenium</groupId> <artifactId>selenium-java</artifactId> <version>4.15.0</version> </dependency> <!-- Jsoup --> <dependency> <groupId>org.jsoup</groupId> <artifactId>jsoup</artifactId> <version>1.17.2</version> </dependency> </dependencies>
Then, use this code to render the page and parse the table:
import org.jsoup.Jsoup; import org.jsoup.nodes.Document; import org.openqa.selenium.WebDriver; import org.openqa.selenium.chrome.ChromeDriver; import org.openqa.selenium.chrome.ChromeOptions; import org.openqa.selenium.support.ui.WebDriverWait; import org.openqa.selenium.support.ui.ExpectedConditions; import java.time.Duration; public class DynamicLeadParser { public static void main(String[] args) { // Configure ChromeDriver (ensure it matches your Chrome browser version) System.setProperty("webdriver.chrome.driver", "path/to/your/chromedriver"); // Use headless mode to avoid opening a visible browser window (optional but great for servers) ChromeOptions options = new ChromeOptions(); options.addArguments("--headless=new"); WebDriver driver = new ChromeDriver(options); WebDriverWait wait = new WebDriverWait(driver, Duration.ofSeconds(15)); try { // Load your target page driver.get("your-target-page-url-here"); // Execute the UpdateDiv() function to trigger the AJAX table load driver.executeScript("UpdateDiv();"); // Wait for the target table to appear (more reliable than hardcoding sleep) wait.until(ExpectedConditions.presenceOfElementLocated(org.openqa.selenium.By.cssSelector("table.none"))); // Grab the fully rendered page source String renderedHtml = driver.getPageSource(); // Now parse this HTML with Jsoup Document doc = Jsoup.parse(renderedHtml); // Extract your lead data from the table var leadTable = doc.selectFirst("table.none"); String fullName = leadTable.select("tr.tm_tt_body td.typedata1[colspan=3]").first().text().trim(); String phoneNumber = leadTable.select("b:contains(222-222-2222)").text().trim(); String email = leadTable.select("#ld_email").attr("href").replace("mailto:", "").split("\\?subject")[0].trim(); // Print or process the extracted data System.out.println("Lead Name: " + fullName); System.out.println("Lead Phone: " + phoneNumber); System.out.println("Lead Email: " + email); } catch (Exception e) { e.printStackTrace(); } finally { // Always quit the driver to clean up resources driver.quit(); } } }
Approach 2: Extract Data Directly with Selenium
Since you're already using Selenium to render the page, you can skip Jsoup entirely and extract the data directly using Selenium's element locators. This is simpler if you don't need Jsoup's specific parsing features.
Example Code Snippet
import org.openqa.selenium.WebDriver; import org.openqa.selenium.chrome.ChromeDriver; import org.openqa.selenium.chrome.ChromeOptions; import org.openqa.selenium.WebElement; import org.openqa.selenium.support.ui.WebDriverWait; import org.openqa.selenium.support.ui.ExpectedConditions; import java.time.Duration; public class SeleniumDirectExtractor { public static void main(String[] args) { System.setProperty("webdriver.chrome.driver", "path/to/your/chromedriver"); ChromeOptions options = new ChromeOptions(); options.addArguments("--headless=new"); WebDriver driver = new ChromeDriver(options); WebDriverWait wait = new WebDriverWait(driver, Duration.ofSeconds(15)); try { driver.get("your-target-page-url-here"); driver.executeScript("UpdateDiv();"); // Wait for the table to load WebElement leadTable = wait.until(ExpectedConditions.presenceOfElementLocated(org.openqa.selenium.By.cssSelector("table.none"))); // Extract data directly via Selenium locators WebElement nameElement = leadTable.findElement(org.openqa.selenium.By.cssSelector("tr.tm_tt_body td.typedata1[colspan=3]")); String fullName = nameElement.getText().trim(); WebElement phoneElement = leadTable.findElement(org.openqa.selenium.By.xpath("//a[contains(text(), '222-222-2222')]")); String phoneNumber = phoneElement.getText().trim(); WebElement emailElement = leadTable.findElement(org.openqa.selenium.By.id("ld_email")); String email = emailElement.getAttribute("href").replace("mailto:", "").split("\\?subject")[0].trim(); System.out.println("Lead Name: " + fullName); System.out.println("Lead Phone: " + phoneNumber); System.out.println("Lead Email: " + email); } catch (Exception e) { e.printStackTrace(); } finally { driver.quit(); } } }
Key Notes
- Browser Driver Compatibility: Make sure your ChromeDriver version matches your installed Chrome browser version (download from ChromeDriver official site—but wait, no external links allowed, so just remember to get the matching version).
- Reliable Waiting: Use
WebDriverWaitwith explicit conditions (like waiting for the table to exist) instead ofThread.sleep()—this makes your code more robust to varying network speeds. - Headless Mode: The
--headless=newflag runs Chrome without a visible window, which is ideal for running scripts on servers or in background processes.
内容的提问来源于stack exchange,提问作者Brandon Woodruff

