Java自动化无法稳定获取Amazon搜索全量结果问题求助
我正在实现一个自动化场景:在amazon.in搜索iPhone,获取所有搜索结果的标题。观察到每页通常有18条结果,但最后一页(第20页)有时结果不足10条。
目前能获取结果,但数量不稳定:有时某页能拿到18条,其他页却不足18条。
尝试过presenceOfAllElementsLocated、visibility等ExpectedConditions等待条件,但均无效;使用numberOfElementsToBeMoreThan能生效,但最后一页结果仍不足10条。另外,使用Thread.sleep可以得到正确结果。使用Java开发,请求解决该问题。
主代码
public static void main(String[] args) throws InterruptedException { browserUtil.openURL("https://amazon.in"); ElementUtils elementUtils = new ElementUtils(driver); By searchField = By.id("twotabsearchtextbox"); By searchButton = By.id("nav-search-submit-button"); By eachResult = By.xpath("//div[starts-with(@cel_widget_id, 'MAIN-SEARCH_RESULTS')]//h2//span"); By nextButton = By.className("s-pagination-next"); elementUtils.sendKeysToTextField(searchField, "iPhone"); elementUtils.clickOnElement(searchButton); System.out.println(elementUtils.getAllElementsTextBySize(eachResult, 17)); while(!Objects.equals(elementUtils.getWebElement(nextButton).getAttribute("aria-disabled"), "true")){ elementUtils.clickOnElement(nextButton); System.out.println(elementUtils.getAllElementsTextBySize(eachResult, 17)); } browserUtil.closeBrowser(); }
ElementUtils工具类代码
public ElementUtils(WebDriver driver){ this.driver = driver; wait = new WebDriverWait(driver, Duration.ofSeconds(20)); actions = new Actions(driver); } public WebElement getWebElement(By locator){ return wait.until(ExpectedConditions.presenceOfElementLocated(locator)); } public List<WebElement> getWebElements(By locator){ return wait.until(ExpectedConditions.presenceOfAllElementsLocatedBy(locator)); } public List<WebElement> getWebElementsBySize(By locator, int size){ return wait.until(ExpectedConditions.numberOfElementsToBeMoreThan(locator, size)); }
1. 区分普通页与最后一页的等待逻辑
当前numberOfElementsToBeMoreThan固定传17,导致最后一页结果不足17条时无法满足等待条件。可以先判断当前页是否为最后一页,再调整等待规则:
- 点击下一页后,先检查
nextButton的aria-disabled属性 - 若为最后一页,使用
numberOfElementsToBeLessThanOrEqualTo(locator, 18)配合元素可见性等待;若为普通页,继续用numberOfElementsToBeMoreThan(locator, 17)
2. 自定义等待条件,确保元素数量稳定
Amazon搜索结果可能存在动态渲染延迟,默认等待条件无法覆盖这种场景。可以自定义等待逻辑,等待元素数量在短时间内保持不变:
public List<WebElement> waitForElementsStable(By locator, int timeoutSeconds) { WebDriverWait wait = new WebDriverWait(driver, Duration.ofSeconds(timeoutSeconds)); return wait.until(driver -> { List<WebElement> elements = driver.findElements(locator); int initialSize = elements.size(); // 等待1秒验证数量是否稳定 try { Thread.sleep(1000); } catch (InterruptedException e) { Thread.currentThread().interrupt(); } List<WebElement> updatedElements = driver.findElements(locator); return updatedElements.size() == initialSize ? updatedElements : null; }); }
调用时替换原有的getAllElementsTextBySize,用这个方法等待元素稳定后再获取文本。
3. 等待页面完全加载
点击下一页后,先等待页面document处于complete状态,避免提前获取未加载完成的元素:
public void waitForPageComplete() { WebDriverWait wait = new WebDriverWait(driver, Duration.ofSeconds(20)); wait.until(driver -> ((JavascriptExecutor) driver).executeScript("return document.readyState").equals("complete")); }
在点击nextButton后,先执行这个方法,再去获取结果元素。
4. 优化元素定位器
当前XPath依赖动态生成的cel_widget_id,稳定性不足。改用基于标准组件的定位器:
By eachResult = By.cssSelector("[data-component-type='s-search-result'] h2 span");
这个定位器依赖Amazon搜索结果的固定组件属性,避免因动态ID变化导致的定位失败。
5. 触发滚动加载
部分结果需要滚动页面才会渲染,在获取元素前先滚动到页面底部:
public void scrollToPageBottom() { ((JavascriptExecutor) driver).executeScript("window.scrollTo(0, document.body.scrollHeight)"); }
执行滚动后再调用等待方法,确保所有结果都被渲染。
内容的提问来源于stack exchange,提问作者stranger

