You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium爬取Web表格时提取Web Elements文本速度过慢,求优化方案

问题:Selenium爬取网页表格速度过慢,求优化方案

我使用Selenium爬取网页表格数据,现有代码会从Web元素(行)列表中读取文本,再将文本存入另一个列表(列),最后调用方法写入Excel。但处理约200行Web元素并写入新列表的过程速度极慢,请问是否有更快的实现方式?还是这种情况属于正常现象?

现有代码

package mypackage;

import java.io.IOException;
import java.time.Duration;
import java.util.ArrayList;
import java.util.List;

import org.openqa.selenium.By;
import org.openqa.selenium.WebDriver;
import org.openqa.selenium.WebElement;
import org.openqa.selenium.chrome.ChromeDriver;

import com.seleniumpractice.utilities.XLUtils;

import io.github.bonigarcia.wdm.WebDriverManager;

public class CovidWebTable {
    static WebDriver driver;
    static XLUtils xl;
    static List<WebElement> header;
    static List<WebElement> rows;
    static List<ArrayList<String>> rowsXL;

    public static void main(String[] args) throws IOException {
        WebDriverManager.chromedriver().setup();
        driver = new ChromeDriver();
        driver.get("https://www.worldometers.info/coronavirus");
        driver.manage().window().maximize();
        driver.manage().timeouts().implicitlyWait(Duration.ofSeconds(10));
        
        WebElement table = driver.findElement(By.xpath("//table[@id='main_table_countries_today']"));
        rows = table.findElements(By.xpath(".//tr[@role='row']"));
        System.out.println("Total rows: "+rows.size());
        
        xl = new XLUtils(".\\datafiles\\covid.xls");
        
        rowsXL = new ArrayList<ArrayList<String>>();
        
        //Add header
        header = table.findElements(By.xpath(".//thead//th"));
        System.out.println("Header cols: "+ header.size());
        
        ArrayList<String> headerXL = new ArrayList<String>();
        
        for(int col=1; col<header.size()-1; col++) {
            headerXL.add(header.get(col).getText());
        }
        
        rowsXL.add(headerXL);
        
        int xlRow = 1;
        int skippedRows = 0;
                
        for(int r=1; r<rows.size(); r++) {
            
            String a = rows.get(r).getText();
            
            //skip empty rows
            if(rows.get(r).getText().equals("")) {
                skippedRows++;
                continue;
            }
            System.out.println("Reading row "+r);   
            
            ArrayList<String> cols = new ArrayList<String>();
            
            for(int c=1; c<header.size(); c++) {
                String data = rows.get(r).findElement(By.xpath(".//td["+(c+1)+"]")).getText();
                cols.add(data);
                
            }
            rowsXL.add(cols);
            xlRow++;
            
        }
        xl.setCellDataFromList(rowsXL, "Orders");
        System.out.println("Scraped Rows: "+ rowsXL.size());
        System.out.println("Skipped Rows: "+skippedRows);
        System.out.println("Complete.");
        
        driver.close();
    }
}

分析与优化方案

处理200行数据速度极慢肯定不属于正常现象,核心问题在于代码中存在大量重复的DOM查找操作,以及不必要的方法调用,以下是具体优化方向:

1. 避免循环内重复查找DOM元素

原代码在每行的循环内,都通过findElement去查找对应td元素,每次查找都会和浏览器通信,开销极大。可以改为一次性获取当前行的所有td元素并缓存:

// 替换原有的行循环逻辑
for(int r=1; r<rows.size(); r++) {
    WebElement currentRow = rows.get(r);
    // 一次性获取当前行所有td
    List<WebElement> cells = currentRow.findElements(By.tagName("td"));
    
    String rowText = currentRow.getText().trim();
    if(rowText.isEmpty()) {
        skippedRows++;
        continue;
    }
    System.out.println("Reading row "+r);   
    
    ArrayList<String> cols = new ArrayList<String>();
    
    for(int c=1; c<header.size(); c++) {
        // 直接使用缓存的cells,无需重复查找
        if(c < cells.size()) {
            cols.add(cells.get(c).getText());
        } else {
            cols.add(""); // 兼容列数不足的情况
        }
    }
    rowsXL.add(cols);
    xlRow++;
}

2. 减少getText()的重复调用

原代码中对同一行两次调用getText(),可以改为只调用一次并缓存结果:

// 原代码:
// String a = rows.get(r).getText();
// if(rows.get(r).getText().equals("")) {

// 优化后:
String rowText = currentRow.getText().trim();
if(rowText.isEmpty()) {
    skippedRows++;
    continue;
}

3. 调整等待策略

隐式等待会作用于所有DOM查找操作,当表格数据已经完全渲染后,不需要再保留长时长的隐式等待,可以在获取表格后关闭或缩短:

WebElement table = driver.findElement(By.xpath("//table[@id='main_table_countries_today']"));
// 获取表格后关闭隐式等待
driver.manage().timeouts().implicitlyWait(Duration.ofSeconds(0));

4. 直接提取HTML解析(大幅提速)

如果表格是静态渲染的,可以直接获取表格的HTML内容,用Jsoup等HTML解析库处理,避免Selenium与浏览器的频繁通信,速度会提升数倍:

// 引入Jsoup依赖后使用
String tableHtml = table.getAttribute("innerHTML");
org.jsoup.nodes.Document doc = org.jsoup.Jsoup.parse(tableHtml);

// 处理表头
Elements headerElements = doc.select("thead th");
ArrayList<String> headerXL = new ArrayList<>();
for(int col=1; col<headerElements.size()-1; col++) {
    headerXL.add(headerElements.get(col).text());
}
rowsXL.add(headerXL);

// 处理数据行
Elements dataRows = doc.select("tr[role='row']");
for(int r=1; r<dataRows.size(); r++) {
    Elements cells = dataRows.get(r).select("td");
    String rowText = dataRows.get(r).text().trim();
    if(rowText.isEmpty()) {
        skippedRows++;
        continue;
    }
    ArrayList<String> cols = new ArrayList<>();
    for(int c=1; c<headerElements.size(); c++) {
        if(c < cells.size()) {
            cols.add(cells.get(c).text());
        } else {
            cols.add("");
        }
    }
    rowsXL.add(cols);
}

5. 确认Excel写入效率

虽然你的代码是先收集所有数据再批量写入,但要确认XLUtils.setCellDataFromList的实现是否高效,比如是否使用了批量写入而非逐单元格写入,如果底层实现低效也会拖慢整体速度。

内容的提问来源于stack exchange,提问作者fdama

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 17:45:28