You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Selenium网页搜索无结果,请求修改代码获取搜索结果URL

解决Google搜索结果URL获取为空的问题

我来帮你修复这个问题,你的代码主要遇到了三个核心问题:Selenium API版本兼容、Google页面结构更新,以及缺少页面加载等待机制。下面是修改后的完整代码,能稳定获取前20条搜索结果的URL:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import TimeoutException

def get_results(search_term, max_results=20):
    url = "https://www.google.com"
    browser = webdriver.Safari()
    browser.get(url)
    
    try:
        # 处理Google的Cookie同意弹窗(如果出现)
        cookie_accept = WebDriverWait(browser, 10).until(
            EC.element_to_be_clickable((By.XPATH, "//div[text()='同意' or text()='Accept all']"))
        )
        cookie_accept.click()
    except TimeoutException:
        # 如果没有弹窗,直接跳过
        pass
    
    # 使用新版Selenium API定位搜索框,等待搜索框加载完成
    search_box = WebDriverWait(browser, 10).until(
        EC.element_to_be_clickable((By.NAME, "q"))
    )
    search_box.send_keys(search_term)
    search_box.submit()
    
    # 等待搜索结果加载完成,定位结果中的链接
    try:
        # 适配当前Google搜索结果的结构:结果项在div#search下的div.g中,链接在h3对应的a标签
        results_container = WebDriverWait(browser, 15).until(
            EC.presence_of_element_located((By.ID, "search"))
        )
        links = results_container.find_elements(By.XPATH, ".//div[@class='g']//h3//a")
    except TimeoutException:
        # 备用定位方式,防止页面结构小幅度变化
        links = browser.find_elements(By.XPATH, "//h3//a")
    
    # 提取URL并限制数量
    results = []
    for link in links[:max_results]:
        href = link.get_attribute("href")
        # 过滤掉Google自身的跳转链接(如果有的话)
        if href and not href.startswith("https://www.google.com/url?"):
            results.append(href)
    
    browser.quit()  # 用quit()比close()更彻底,释放资源
    return results

# 测试调用
print(get_results("fish", max_results=10))

关键改动说明:

  • 更新Selenium定位API:替换了已弃用的find_element_by_name等方法,改用By类配合WebDriverWait的新版API,兼容性更好。
  • 添加显式等待:所有元素定位都加入了等待,确保页面加载完成后再操作,避免因元素未加载导致的空列表。
  • 适配Google页面结构:当前Google搜索结果的容器是div#search下的div.g,调整了xpath选择器来匹配最新的页面结构。
  • 处理Cookie弹窗:自动处理常见的Cookie同意弹窗,避免弹窗阻挡后续操作。
  • 结果数量限制:通过links[:max_results]轻松控制返回结果的数量,默认20条,可自定义。
  • 过滤跳转链接:排除Google的中转跳转链接,直接返回目标网站的真实URL。

内容的提问来源于stack exchange,提问作者Baobab1988

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 09:30:33