You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python发起谷歌搜索请求?现有库失效求解决方案

如何用Python发起谷歌搜索请求

你的Selenium代码问题排查与修复

你的代码无法获取搜索结果,核心原因有两个:谷歌的反爬机制会识别无头浏览器环境,以及你使用的CSS选择器(如div[data-sokoban-container])已失效——谷歌会频繁更新搜索页面的DOM结构。

以下是修复后的Selenium代码,解决了上述问题:

from pathlib import Path
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.firefox.options import Options as FirefoxOptions
from selenium.webdriver.firefox.service import Service as FirefoxService
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

QUERY = input("输入谷歌搜索关键词: ")

base_dir = Path(__file__).resolve().parent
driver_path = base_dir.parent / "geckodriver.exe"

def run_with_selenium() -> bool:
    options = FirefoxOptions()
    options.add_argument("--headless")
    # 添加模拟真实浏览器的参数,规避反爬识别
    options.add_argument("--window-size=1920,1080")
    options.add_argument("--disable-blink-features=AutomationControlled")
    options.add_argument("--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36")
    options.add_experimental_option("excludeSwitches", ["enable-automation"])
    options.add_experimental_option("useAutomationExtension", False)

    service = FirefoxService(executable_path=str(driver_path))
    driver = None
    try:
        driver = webdriver.Firefox(service=service, options=options)
        # 移除Selenium自动化标识
        driver.execute_script("Object.defineProperty(navigator, 'webdriver', {get: () => undefined})")
        
        driver.get(f"https://www.google.com/search?q={QUERY}")
        
        # 用显式等待替代time.sleep,确保元素加载完成
        wait = WebDriverWait(driver, 10)
        # 处理谷歌Cookie同意弹窗
        try:
            accept_btn = wait.until(EC.element_to_be_clickable((By.XPATH, "//div[text()='同意']")))
            accept_btn.click()
        except:
            pass
        
        # 使用当前有效的搜索结果容器选择器
        result_containers = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "div.g")))
        
        print(f"\n'{QUERY}' 的搜索结果:\n")
        results = []
        
        for idx, container in enumerate(result_containers[:10], 1):
            try:
                title_elem = container.find_element(By.CSS_SELECTOR, "h3")
                title = title_elem.text
                
                link_elem = container.find_element(By.CSS_SELECTOR, "a")
                link = link_elem.get_attribute("href")
                
                snippet_elem = container.find_element(By.CSS_SELECTOR, "div.VwiC3b")
                snippet = snippet_elem.text
                
                if title and link:
                    print(f"{idx}. {title}")
                    print(f"   链接: {link}")
                    if snippet:
                        print(f"   摘要: {snippet}")
                    print()
                    results.append({"title": title, "url": link, "snippet": snippet})
            except Exception:
                continue
        
        if results:
            print(f"✓ 找到 {len(results)} 条结果")
            return True
        else:
            print("未提取到结果,但页面加载成功")
            print(f"页面标题: {driver.title}")
            return True
            
    except Exception as error:
        print(f"Selenium错误: {type(error).__name__}: {str(error)[:200]}")
        return False
    finally:
        if driver:
            driver.quit()

run_with_selenium()

关键改动说明

  • 新增多个浏览器参数,让无头模式更接近真实用户环境,避免被反爬拦截
  • 替换失效的CSS选择器,使用当前谷歌搜索结果的标准容器选择器div.g和摘要选择器div.VwiC3b
  • 用WebDriverWait显式等待替代time.sleep,确保元素加载完成后再提取
  • 处理谷歌Cookie同意弹窗,避免弹窗阻塞后续操作
  • 禁用Selenium的webdriver标识,进一步规避反爬检测

其他可用的Python库

1. googlesearch-python

PyPi上的轻量级库,封装了谷歌搜索逻辑,无需自行处理页面解析和基础反爬配置:

from googlesearch import search

QUERY = input("输入谷歌搜索关键词: ")

# 获取前10条结果
for idx, result in enumerate(search(QUERY, num_results=10), 1):
    print(f"{idx}. {result}")

2. requests + BeautifulSoup(需自行处理反爬)

如果不想用Selenium或第三方库,可直接用requests发送请求配合BeautifulSoup解析页面,但需配置合适的请求头,高频率请求可能触发谷歌人机验证:

import requests
from bs4 import BeautifulSoup

QUERY = input("输入谷歌搜索关键词: ")
URL = f"https://www.google.com/search?q={QUERY}"

headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
}

response = requests.get(URL, headers=headers)
soup = BeautifulSoup(response.text, "html.parser")

results = soup.find_all("div", class_="g")
for idx, result in enumerate(results[:10], 1):
    try:
        title = result.find("h3").text
        link = result.find("a")["href"]
        snippet = result.find("div", class_="VwiC3b").text
        print(f"{idx}. {title}")
        print(f"   链接: {link}")
        print(f"   摘要: {snippet}\n")
    except:
        continue

内容的提问来源于stack exchange,提问作者Vedika Gupta

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.02 01:03:10