You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Selenium+Java批量提取HTML中所有文本及链接属性

批量提取分类文本与对应链接的Java Selenium解决方案

方法一:直接定位目标<a>元素集合

直接定位包含分类标题的<a>标签,一次性提取文本和链接属性:

// 定位所有分类对应的<a>元素
List<WebElement> categoryLinks = driver.findElements(By.cssSelector("div.col-12 > a"));

// 遍历集合提取信息
for (WebElement link : categoryLinks) {
    // 获取分类标题文本(去除首尾空格)
    String categoryText = link.findElement(By.className("category_title")).getText().trim();
    // 获取链接地址
    String categoryHref = link.getAttribute("href");
    // 按需输出或存储数据
    System.out.printf("分类:%s,链接:%s%n", categoryText, categoryHref);
}

方法二:通过已定位的h5元素反向获取父级链接

如果你已经能拿到h5元素列表,可通过相对定位找到其父级<a>标签:

// 定位所有分类标题h5元素
List<WebElement> categoryTitles = driver.findElements(By.className("category_title"));

// 遍历每个标题,获取对应链接
for (WebElement title : categoryTitles) {
    String categoryText = title.getText().trim();
    // 通过XPath查找最近的父级<a>元素
    WebElement parentLink = title.findElement(By.xpath("./ancestor::a[1]"));
    String categoryHref = parentLink.getAttribute("href");
    // 按需输出或存储数据
    System.out.printf("分类:%s,链接:%s%n", categoryText, categoryHref);
}

额外注意事项

  • 确保页面元素加载完成,建议添加显式等待避免NoSuchElementException:
WebDriverWait wait = new WebDriverWait(driver, Duration.ofSeconds(10));
List<WebElement> categoryLinks = wait.until(
    ExpectedConditions.visibilityOfAllElementsLocatedBy(By.cssSelector("div.col-12 > a"))
);
  • 如果获取的是相对路径,可拼接为绝对路径:
String baseUrl = driver.getCurrentUrl();
String absoluteHref = new URL(new URL(baseUrl), categoryHref).toString();

内容的提问来源于stack exchange,提问作者Kevin Grant

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 17:35:25