使用Selenium无法获取Google新闻页面中的新闻文章链接
解决Google新闻文章链接无法获取的问题
嘿,我来帮你搞定这个问题!你遇到的核心问题有两个:一是页面加载未完成就开始抓取元素,导致动态加载的新闻链接没被捕获;二是定位范围太宽泛,拿到了页面上所有的a标签(比如导航、侧边栏的链接),却没精准定位到新闻文章的链接。
问题分析
- 动态加载与等待缺失:点击「新闻」按钮后,Google新闻的内容是动态渲染的,直接调用
findElements会漏掉还没加载出来的新闻链接。 - 定位选择器不准确:
By.tagName("a")会匹配页面上所有的超链接,但新闻文章的链接通常嵌套在特定的容器元素里(比如Google新闻里的文章卡片),需要更精准的选择器来过滤。
修复后的代码
import org.openqa.selenium.By; import org.openqa.selenium.WebDriver; import org.openqa.selenium.WebElement; import org.openqa.selenium.firefox.FirefoxDriver; import org.openqa.selenium.support.ui.ExpectedConditions; import org.openqa.selenium.support.ui.WebDriverWait; import java.io.BufferedWriter; import java.io.FileWriter; import java.io.IOException; import java.util.List; public class GoogleNewsScraper { public static void main(String[] args) throws IOException { String baseURL = "https://www.google.com"; WebDriver driver = new FirefoxDriver(); WebDriverWait wait = new WebDriverWait(driver, 10); // 10秒显式等待 try { driver.manage().window().maximize(); driver.get(baseURL); // 等待搜索框加载并输入关键词 WebElement searchBar = wait.until(ExpectedConditions.visibilityOfElementLocated(By.xpath("//*[@id='lst-ib']"))); searchBar.sendKeys("Liverpool"); // 等待搜索按钮可点击并点击 WebElement searchButton = wait.until(ExpectedConditions.elementToBeClickable(By.name("btnK"))); searchButton.click(); // 等待新闻标签可点击并切换到新闻页面 WebElement newsTab = wait.until(ExpectedConditions.elementToBeClickable(By.xpath("//*[@id='hdtb-msb-vis']/div[2]/a"))); newsTab.click(); // 等待新闻文章容器加载完成,定位所有新闻文章的链接 // 这里用css选择器定位新闻卡片内的a标签,适配Google新闻的页面结构 List<WebElement> newsLinks = wait.until(ExpectedConditions.presenceOfAllElementsLocatedBy( By.cssSelector("div.g a[href^='https://']") )); System.out.println("找到的新闻链接数量:" + newsLinks.size()); // 写入文件并打印到控制台 try (FileWriter file = new FileWriter("/Users/lekharaj/Desktop/LFC.txt"); BufferedWriter writer = new BufferedWriter(file)) { for (WebElement link : newsLinks) { String href = link.getAttribute("href"); // 过滤掉Google的跳转链接(如果有的话) if (href != null && !href.startsWith("https://www.google.com/url?")) { System.out.println(href); writer.write(href); writer.newLine(); } } } } finally { driver.quit(); // 关闭浏览器,释放资源 } } }
关键改动说明
- 添加显式等待:用
WebDriverWait替代直接查找元素,确保目标元素加载完成后再操作,避免动态加载导致的元素缺失。 - 精准定位新闻链接:使用
By.cssSelector("div.g a[href^='https://']")定位新闻卡片(class为g的div)内的外部链接,过滤掉Google内部的跳转链接。 - 资源自动管理:用try-with-resources语法自动关闭文件流,避免资源泄漏。
- 关闭浏览器:在finally块中调用
driver.quit(),确保程序结束后关闭浏览器进程。
额外提示
如果Google新闻的页面结构后续发生变化,你可以通过浏览器的开发者工具(F12)查看新闻文章链接的父容器class或属性,调整选择器即可。比如有时候新闻卡片的class可能是dbsr,那选择器可以改成By.cssSelector("div.dbsr a")。
内容的提问来源于stack exchange,提问作者harsha ray
相关产品推荐
相关产品推荐

