如何用Selenium+Java批量提取HTML中所有文本及链接属性
批量提取分类文本与对应链接的Java Selenium解决方案
方法一:直接定位目标<a>元素集合
直接定位包含分类标题的<a>标签,一次性提取文本和链接属性:
// 定位所有分类对应的<a>元素 List<WebElement> categoryLinks = driver.findElements(By.cssSelector("div.col-12 > a")); // 遍历集合提取信息 for (WebElement link : categoryLinks) { // 获取分类标题文本(去除首尾空格) String categoryText = link.findElement(By.className("category_title")).getText().trim(); // 获取链接地址 String categoryHref = link.getAttribute("href"); // 按需输出或存储数据 System.out.printf("分类:%s,链接:%s%n", categoryText, categoryHref); }
方法二:通过已定位的h5元素反向获取父级链接
如果你已经能拿到h5元素列表,可通过相对定位找到其父级<a>标签:
// 定位所有分类标题h5元素 List<WebElement> categoryTitles = driver.findElements(By.className("category_title")); // 遍历每个标题,获取对应链接 for (WebElement title : categoryTitles) { String categoryText = title.getText().trim(); // 通过XPath查找最近的父级<a>元素 WebElement parentLink = title.findElement(By.xpath("./ancestor::a[1]")); String categoryHref = parentLink.getAttribute("href"); // 按需输出或存储数据 System.out.printf("分类:%s,链接:%s%n", categoryText, categoryHref); }
额外注意事项
- 确保页面元素加载完成,建议添加显式等待避免
NoSuchElementException:
WebDriverWait wait = new WebDriverWait(driver, Duration.ofSeconds(10)); List<WebElement> categoryLinks = wait.until( ExpectedConditions.visibilityOfAllElementsLocatedBy(By.cssSelector("div.col-12 > a")) );
- 如果获取的是相对路径,可拼接为绝对路径:
String baseUrl = driver.getCurrentUrl(); String absoluteHref = new URL(new URL(baseUrl), categoryHref).toString();
内容的提问来源于stack exchange,提问作者Kevin Grant
相关产品推荐
相关产品推荐

