You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium无法获取Google新闻页面中的新闻文章链接

解决Google新闻文章链接无法获取的问题

嘿,我来帮你搞定这个问题!你遇到的核心问题有两个:一是页面加载未完成就开始抓取元素,导致动态加载的新闻链接没被捕获;二是定位范围太宽泛,拿到了页面上所有的a标签(比如导航、侧边栏的链接),却没精准定位到新闻文章的链接。

问题分析

  1. 动态加载与等待缺失:点击「新闻」按钮后,Google新闻的内容是动态渲染的,直接调用findElements会漏掉还没加载出来的新闻链接。
  2. 定位选择器不准确:By.tagName("a")会匹配页面上所有的超链接,但新闻文章的链接通常嵌套在特定的容器元素里(比如Google新闻里的文章卡片),需要更精准的选择器来过滤。

修复后的代码

import org.openqa.selenium.By;
import org.openqa.selenium.WebDriver;
import org.openqa.selenium.WebElement;
import org.openqa.selenium.firefox.FirefoxDriver;
import org.openqa.selenium.support.ui.ExpectedConditions;
import org.openqa.selenium.support.ui.WebDriverWait;

import java.io.BufferedWriter;
import java.io.FileWriter;
import java.io.IOException;
import java.util.List;

public class GoogleNewsScraper {
    public static void main(String[] args) throws IOException {
        String baseURL = "https://www.google.com";
        WebDriver driver = new FirefoxDriver();
        WebDriverWait wait = new WebDriverWait(driver, 10); // 10秒显式等待

        try {
            driver.manage().window().maximize();
            driver.get(baseURL);

            // 等待搜索框加载并输入关键词
            WebElement searchBar = wait.until(ExpectedConditions.visibilityOfElementLocated(By.xpath("//*[@id='lst-ib']")));
            searchBar.sendKeys("Liverpool");

            // 等待搜索按钮可点击并点击
            WebElement searchButton = wait.until(ExpectedConditions.elementToBeClickable(By.name("btnK")));
            searchButton.click();

            // 等待新闻标签可点击并切换到新闻页面
            WebElement newsTab = wait.until(ExpectedConditions.elementToBeClickable(By.xpath("//*[@id='hdtb-msb-vis']/div[2]/a")));
            newsTab.click();

            // 等待新闻文章容器加载完成,定位所有新闻文章的链接
            // 这里用css选择器定位新闻卡片内的a标签,适配Google新闻的页面结构
            List<WebElement> newsLinks = wait.until(ExpectedConditions.presenceOfAllElementsLocatedBy(
                    By.cssSelector("div.g a[href^='https://']")
            ));

            System.out.println("找到的新闻链接数量:" + newsLinks.size());

            // 写入文件并打印到控制台
            try (FileWriter file = new FileWriter("/Users/lekharaj/Desktop/LFC.txt");
                 BufferedWriter writer = new BufferedWriter(file)) {

                for (WebElement link : newsLinks) {
                    String href = link.getAttribute("href");
                    // 过滤掉Google的跳转链接(如果有的话)
                    if (href != null && !href.startsWith("https://www.google.com/url?")) {
                        System.out.println(href);
                        writer.write(href);
                        writer.newLine();
                    }
                }
            }
        } finally {
            driver.quit(); // 关闭浏览器,释放资源
        }
    }
}

关键改动说明

  • 添加显式等待:用WebDriverWait替代直接查找元素,确保目标元素加载完成后再操作,避免动态加载导致的元素缺失。
  • 精准定位新闻链接:使用By.cssSelector("div.g a[href^='https://']")定位新闻卡片(class为g的div)内的外部链接,过滤掉Google内部的跳转链接。
  • 资源自动管理:用try-with-resources语法自动关闭文件流,避免资源泄漏。
  • 关闭浏览器:在finally块中调用driver.quit(),确保程序结束后关闭浏览器进程。

额外提示

如果Google新闻的页面结构后续发生变化,你可以通过浏览器的开发者工具(F12)查看新闻文章链接的父容器class或属性,调整选择器即可。比如有时候新闻卡片的class可能是dbsr,那选择器可以改成By.cssSelector("div.dbsr a")。

内容的提问来源于stack exchange,提问作者harsha ray

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 09:25:37