在Eclipse中使用Java抓取JS渲染网站,提取表格百分比的技术问询
Hey there! I totally get the frustration here—dealing with JavaScript-rendered tables in Java can feel like hitting a wall when the raw HTML doesn’t have the percentage data you need. Let’s break down the most reliable ways to solve this, so you can grab those "61% | 25% | 14%" strings and other percentage values successfully.
This approach mimics a real browser, so it fully renders the page after JavaScript runs—perfect for tables that are built client-side.
Step 1: Add Dependencies (Maven Example)
First, include the Selenium and ChromeDriver dependencies in your pom.xml:
<dependency> <groupId>org.seleniumhq.selenium</groupId> <artifactId>selenium-java</artifactId> <version>4.15.0</version> </dependency> <dependency> <groupId>org.seleniumhq.selenium</groupId> <artifactId>selenium-chrome-driver</artifactId> <version>4.15.0</version> </dependency>
Step 2: Write the Scraping Code
This code will launch a headless Chrome instance, load the page, wait for the table to render, and extract the percentage values:
import org.openqa.selenium.By; import org.openqa.selenium.WebDriver; import org.openqa.selenium.WebElement; import org.openqa.selenium.chrome.ChromeDriver; import org.openqa.selenium.chrome.ChromeOptions; import org.openqa.selenium.support.ui.WebDriverWait; import org.openqa.selenium.support.ui.ExpectedConditions; import java.time.Duration; import java.util.List; public class JsTableScraper { public static void main(String[] args) { // Configure headless Chrome ChromeOptions options = new ChromeOptions(); options.addArguments("--headless=new"); // Modern headless mode options.addArguments("--disable-gpu"); options.addArguments("--no-sandbox"); // Initialize driver (try-with-resources ensures it closes properly) try (WebDriver driver = new ChromeDriver(options)) { driver.get("YOUR_TARGET_URL"); // Wait for the percentage elements to load (way better than hardcoding sleep) WebDriverWait wait = new WebDriverWait(driver, Duration.ofSeconds(10)); wait.until(ExpectedConditions.presenceOfElementLocated(By.xpath("//td[contains(text(), '%')]"))); // Grab all elements containing percentages List<WebElement> percentageCells = driver.findElements(By.xpath("//td[contains(text(), '%')]")); // Extract and print the values for (WebElement cell : percentageCells) { String value = cell.getText().trim(); System.out.println(value); // If you need to group values by table row, loop through <tr> elements first } } catch (Exception e) { e.printStackTrace(); } } }
If you don’t want to rely on an external browser, HtmlUnit is a pure-Java headless browser that can execute JavaScript. It’s lighter than Selenium but might struggle with super complex JS frameworks.
Step 1: Add Dependency (Maven)
<dependency> <groupId>net.sourceforge.htmlunit</groupId> <artifactId>htmlunit</artifactId> <version>2.70.0</version> </dependency>
Step 2: Write the Scraping Code
import com.gargoylesoftware.htmlunit.WebClient; import com.gargoylesoftware.htmlunit.html.HtmlPage; import com.gargoylesoftware.htmlunit.html.HtmlTable; import com.gargoylesoftware.htmlunit.html.HtmlTableCell; import com.gargoylesoftware.htmlunit.html.HtmlTableRow; import java.io.IOException; public class HtmlUnitTableScraper { public static void main(String[] args) { try (WebClient webClient = new WebClient()) { // Enable JavaScript and tweak settings for speed webClient.getOptions().setJavaScriptEnabled(true); webClient.getOptions().setCssEnabled(false); // Skip CSS rendering to save time webClient.getOptions().setThrowExceptionOnScriptError(false); // Ignore minor JS errors // Load the page and wait for background JS to finish HtmlPage page = webClient.getPage("YOUR_TARGET_URL"); webClient.waitForBackgroundJavaScript(5000); // Wait up to 5 seconds for JS // Locate the table (use ID, class, or XPath depending on the page) HtmlTable targetTable = page.getHtmlElementById("TABLE_ID"); // Alternative: HtmlTable targetTable = (HtmlTable) page.getByXPath("//table[@class='table-class-name']").get(0); // Loop through rows and cells to find percentages for (HtmlTableRow row : targetTable.getRows()) { for (HtmlTableCell cell : row.getCells()) { String cellText = cell.getTextContent().trim(); if (cellText.contains("%")) { System.out.println(cellText); } } } } catch (IOException e) { e.printStackTrace(); } } }
Before diving into browser rendering, check if the page pulls the table data from an API. Use your browser’s DevTools → Network tab to look for XHR/fetch requests that return JSON with the percentage values. If you find one, you can use Java’s HttpClient or OkHttp to call the API directly—this is faster and more reliable than rendering the whole page.
Quick Example with HttpClient
import java.net.URI; import java.net.http.HttpClient; import java.net.http.HttpRequest; import java.net.http.HttpResponse; public class ApiScraper { public static void main(String[] args) throws Exception { HttpClient client = HttpClient.newHttpClient(); HttpRequest request = HttpRequest.newBuilder() .uri(URI.create("THE_API_ENDPOINT_URL")) .header("User-Agent", "Mozilla/5.0") // Match browser UA to avoid blocks .build(); HttpResponse<String> response = client.send(request, HttpResponse.BodyHandlers.ofString()); // Parse the JSON response with Jackson or Gson to extract percentages System.out.println(response.body()); } }
Just remember to watch out for anti-scraping measures: set a valid User-Agent, respect robots.txt, and don’t hammer the site with too many requests too quickly.
内容的提问来源于stack exchange,提问作者Javier

