如何用Python发起谷歌搜索请求?现有库失效求解决方案
如何用Python发起谷歌搜索请求
你的Selenium代码问题排查与修复
你的代码无法获取搜索结果,核心原因有两个:谷歌的反爬机制会识别无头浏览器环境,以及你使用的CSS选择器(如div[data-sokoban-container])已失效——谷歌会频繁更新搜索页面的DOM结构。
以下是修复后的Selenium代码,解决了上述问题:
from pathlib import Path from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.firefox.options import Options as FirefoxOptions from selenium.webdriver.firefox.service import Service as FirefoxService from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC QUERY = input("输入谷歌搜索关键词: ") base_dir = Path(__file__).resolve().parent driver_path = base_dir.parent / "geckodriver.exe" def run_with_selenium() -> bool: options = FirefoxOptions() options.add_argument("--headless") # 添加模拟真实浏览器的参数,规避反爬识别 options.add_argument("--window-size=1920,1080") options.add_argument("--disable-blink-features=AutomationControlled") options.add_argument("--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36") options.add_experimental_option("excludeSwitches", ["enable-automation"]) options.add_experimental_option("useAutomationExtension", False) service = FirefoxService(executable_path=str(driver_path)) driver = None try: driver = webdriver.Firefox(service=service, options=options) # 移除Selenium自动化标识 driver.execute_script("Object.defineProperty(navigator, 'webdriver', {get: () => undefined})") driver.get(f"https://www.google.com/search?q={QUERY}") # 用显式等待替代time.sleep,确保元素加载完成 wait = WebDriverWait(driver, 10) # 处理谷歌Cookie同意弹窗 try: accept_btn = wait.until(EC.element_to_be_clickable((By.XPATH, "//div[text()='同意']"))) accept_btn.click() except: pass # 使用当前有效的搜索结果容器选择器 result_containers = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "div.g"))) print(f"\n'{QUERY}' 的搜索结果:\n") results = [] for idx, container in enumerate(result_containers[:10], 1): try: title_elem = container.find_element(By.CSS_SELECTOR, "h3") title = title_elem.text link_elem = container.find_element(By.CSS_SELECTOR, "a") link = link_elem.get_attribute("href") snippet_elem = container.find_element(By.CSS_SELECTOR, "div.VwiC3b") snippet = snippet_elem.text if title and link: print(f"{idx}. {title}") print(f" 链接: {link}") if snippet: print(f" 摘要: {snippet}") print() results.append({"title": title, "url": link, "snippet": snippet}) except Exception: continue if results: print(f"✓ 找到 {len(results)} 条结果") return True else: print("未提取到结果,但页面加载成功") print(f"页面标题: {driver.title}") return True except Exception as error: print(f"Selenium错误: {type(error).__name__}: {str(error)[:200]}") return False finally: if driver: driver.quit() run_with_selenium()
关键改动说明
- 新增多个浏览器参数,让无头模式更接近真实用户环境,避免被反爬拦截
- 替换失效的CSS选择器,使用当前谷歌搜索结果的标准容器选择器
div.g和摘要选择器div.VwiC3b - 用
WebDriverWait显式等待替代time.sleep,确保元素加载完成后再提取 - 处理谷歌Cookie同意弹窗,避免弹窗阻塞后续操作
- 禁用Selenium的
webdriver标识,进一步规避反爬检测
其他可用的Python库
1. googlesearch-python
PyPi上的轻量级库,封装了谷歌搜索逻辑,无需自行处理页面解析和基础反爬配置:
from googlesearch import search QUERY = input("输入谷歌搜索关键词: ") # 获取前10条结果 for idx, result in enumerate(search(QUERY, num_results=10), 1): print(f"{idx}. {result}")
2. requests + BeautifulSoup(需自行处理反爬)
如果不想用Selenium或第三方库,可直接用requests发送请求配合BeautifulSoup解析页面,但需配置合适的请求头,高频率请求可能触发谷歌人机验证:
import requests from bs4 import BeautifulSoup QUERY = input("输入谷歌搜索关键词: ") URL = f"https://www.google.com/search?q={QUERY}" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } response = requests.get(URL, headers=headers) soup = BeautifulSoup(response.text, "html.parser") results = soup.find_all("div", class_="g") for idx, result in enumerate(results[:10], 1): try: title = result.find("h3").text link = result.find("a")["href"] snippet = result.find("div", class_="VwiC3b").text print(f"{idx}. {title}") print(f" 链接: {link}") print(f" 摘要: {snippet}\n") except: continue
内容的提问来源于stack exchange,提问作者Vedika Gupta
相关产品推荐
相关产品推荐

