You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何获取Google News点击后跳转文章的URL而非原站点URL?

解决Google News点击文章后无法获取真实URL的问题

你这段代码的问题在于:Google News的文章链接点击后,要么是在新标签页打开真实文章,要么是页面通过前端技术伪装(地址栏保留news.google.com域名,实际内容嵌入在iframe里),所以直接取driver.current_url只能拿到Google News的地址,而非目标文章的真实URL。

方案1:直接提取真实链接(推荐,更高效)

Google News的文章标签里已经存储了真实链接,放在data-url属性中,不用点击页面就能直接获取:

# 导航到Google News分类页
driver.get(f"https://news.google.com/topics/{random_category}")
time.sleep(5)

# 获取第一篇文章的元素
article_link = wait.until(EC.visibility_of_element_located((By.CSS_SELECTOR, "article > div > a")))
print(f"标题: {article_link.get_attribute('aria-label')}")
# 直接提取真实链接
article_url = article_link.get_attribute('data-url')
print(f"真实文章URL: {article_url}")

# 若需要打开文章,再执行点击操作
article_link.click()
# 点击后若打开新标签页,需切换到新标签页
driver.switch_to.window(driver.window_handles[-1])
# 此时获取的current_url就是真实地址
print(f"跳转后URL: {driver.current_url}")

方案2:处理点击后的页面跳转(适用于必须打开页面的场景)

如果点击后文章在新标签页打开,需要先切换到新标签页再获取URL;如果内容在iframe里加载,则需切换到iframe再取地址:

# 导航到Google News分类页
driver.get(f"https://news.google.com/topics/{random_category}")
time.sleep(5)

# 获取第一篇文章并点击
article_link = wait.until(EC.visibility_of_element_located((By.CSS_SELECTOR, "article > div > a")))
print(f"标题: {article_link.get_attribute('aria-label')}")
# 记录当前窗口句柄
original_window = driver.current_window_handle
article_link.click()

# 等待新标签页打开并切换到新标签页
wait.until(EC.number_of_windows_to_be(2))
for window_handle in driver.window_handles:
    if window_handle != original_window:
        driver.switch_to.window(window_handle)
        break

# 等待页面加载完成后获取真实URL
wait.until(EC.url_contains("http"))
article_url = driver.current_url
print(f"真实文章URL: {article_url}")

# 若需要切回原Google News页面,可执行以下代码
# driver.switch_to.window(original_window)

内容的提问来源于stack exchange,提问作者Jalen Pless

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 22:04:59