You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java自动化无法稳定获取Amazon搜索全量结果问题求助

问题描述

我正在实现一个自动化场景:在amazon.in搜索iPhone,获取所有搜索结果的标题。观察到每页通常有18条结果,但最后一页(第20页)有时结果不足10条。

目前能获取结果,但数量不稳定:有时某页能拿到18条,其他页却不足18条。

尝试过presenceOfAllElementsLocated、visibility等ExpectedConditions等待条件,但均无效;使用numberOfElementsToBeMoreThan能生效,但最后一页结果仍不足10条。另外,使用Thread.sleep可以得到正确结果。使用Java开发,请求解决该问题。

主代码

public static void main(String[] args) throws InterruptedException {

        browserUtil.openURL("https://amazon.in");
        ElementUtils elementUtils = new ElementUtils(driver);

        By searchField = By.id("twotabsearchtextbox");
        By searchButton = By.id("nav-search-submit-button");
        By eachResult = By.xpath("//div[starts-with(@cel_widget_id, 'MAIN-SEARCH_RESULTS')]//h2//span");
        By nextButton = By.className("s-pagination-next");

        elementUtils.sendKeysToTextField(searchField, "iPhone");
        elementUtils.clickOnElement(searchButton);
        System.out.println(elementUtils.getAllElementsTextBySize(eachResult, 17));

        while(!Objects.equals(elementUtils.getWebElement(nextButton).getAttribute("aria-disabled"), "true")){
            elementUtils.clickOnElement(nextButton);
            System.out.println(elementUtils.getAllElementsTextBySize(eachResult, 17));
        }

        browserUtil.closeBrowser();
}

ElementUtils工具类代码

public ElementUtils(WebDriver driver){
        this.driver = driver;
        wait = new WebDriverWait(driver, Duration.ofSeconds(20));
        actions = new Actions(driver);
    }

    public WebElement getWebElement(By locator){
        return wait.until(ExpectedConditions.presenceOfElementLocated(locator));
    }

    public List<WebElement> getWebElements(By locator){
        return wait.until(ExpectedConditions.presenceOfAllElementsLocatedBy(locator));
    }
    public List<WebElement> getWebElementsBySize(By locator, int size){
        return wait.until(ExpectedConditions.numberOfElementsToBeMoreThan(locator, size));
}
解决方案

1. 区分普通页与最后一页的等待逻辑

当前numberOfElementsToBeMoreThan固定传17,导致最后一页结果不足17条时无法满足等待条件。可以先判断当前页是否为最后一页,再调整等待规则:

  • 点击下一页后,先检查nextButton的aria-disabled属性
  • 若为最后一页,使用numberOfElementsToBeLessThanOrEqualTo(locator, 18)配合元素可见性等待;若为普通页,继续用numberOfElementsToBeMoreThan(locator, 17)

2. 自定义等待条件,确保元素数量稳定

Amazon搜索结果可能存在动态渲染延迟,默认等待条件无法覆盖这种场景。可以自定义等待逻辑,等待元素数量在短时间内保持不变:

public List<WebElement> waitForElementsStable(By locator, int timeoutSeconds) {
    WebDriverWait wait = new WebDriverWait(driver, Duration.ofSeconds(timeoutSeconds));
    return wait.until(driver -> {
        List<WebElement> elements = driver.findElements(locator);
        int initialSize = elements.size();
        // 等待1秒验证数量是否稳定
        try {
            Thread.sleep(1000);
        } catch (InterruptedException e) {
            Thread.currentThread().interrupt();
        }
        List<WebElement> updatedElements = driver.findElements(locator);
        return updatedElements.size() == initialSize ? updatedElements : null;
    });
}

调用时替换原有的getAllElementsTextBySize,用这个方法等待元素稳定后再获取文本。

3. 等待页面完全加载

点击下一页后,先等待页面document处于complete状态,避免提前获取未加载完成的元素:

public void waitForPageComplete() {
    WebDriverWait wait = new WebDriverWait(driver, Duration.ofSeconds(20));
    wait.until(driver -> ((JavascriptExecutor) driver).executeScript("return document.readyState").equals("complete"));
}

在点击nextButton后,先执行这个方法,再去获取结果元素。

4. 优化元素定位器

当前XPath依赖动态生成的cel_widget_id,稳定性不足。改用基于标准组件的定位器:

By eachResult = By.cssSelector("[data-component-type='s-search-result'] h2 span");

这个定位器依赖Amazon搜索结果的固定组件属性,避免因动态ID变化导致的定位失败。

5. 触发滚动加载

部分结果需要滚动页面才会渲染,在获取元素前先滚动到页面底部:

public void scrollToPageBottom() {
    ((JavascriptExecutor) driver).executeScript("window.scrollTo(0, document.body.scrollHeight)");
}

执行滚动后再调用等待方法,确保所有结果都被渲染。

内容的提问来源于stack exchange,提问作者stranger

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 17:25:03