You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python获取谷歌搜索页面HTML代码?求有效解决方法

解决谷歌搜索页面源码提取及元素定位问题

问题根源

谷歌对非浏览器请求有反爬机制,直接用requests.get()不带任何请求头,会返回反爬验证页面而非真实搜索结果,自然找不到目标元素。即使使用selenium,若未配置浏览器参数模拟正常用户行为,也可能被识别为自动化工具,导致页面加载异常。

解决方案

1. 优化requests请求(添加模拟浏览器请求头)

给请求带上浏览器的User-Agent,让谷歌判定为正常用户访问:

import requests
from bs4 import BeautifulSoup

url = "https://www.google.com/search?q=how+to+get+google+search+page+source+code+by+python"
# 替换为你自己浏览器的User-Agent,可在浏览器开发者工具Network面板查看
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}

resp = requests.get(url, headers=headers)
resp.encoding = 'utf-8'
soup = BeautifulSoup(resp.text, 'html.parser')

# 尝试查找目标容器,若yuRUbf失效,用备选方案定位结果链接
search_results = soup.find_all('div', class_="yuRUbf")
if not search_results:
    # 备选:过滤出真实结果链接,排除谷歌自身跳转链接
    all_links = soup.find_all('a', href=True)
    filtered_results = [a for a in all_links if not a['href'].startswith('/') and 'google.com' not in a['href']]
    print(filtered_results)
else:
    print(search_results)

2. 正确使用Selenium(模拟真实浏览器行为)

配置Chrome选项禁用自动化检测,同时等待页面完全加载:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# 配置Chrome参数,模拟正常用户
chrome_options = Options()
chrome_options.add_argument("start-maximized")
chrome_options.add_argument("--disable-blink-features=AutomationControlled")
chrome_options.add_experimental_option("excludeSwitches", ["enable-automation"])
chrome_options.add_experimental_option('useAutomationExtension', False)
chrome_options.add_argument("User-Agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36")

driver = webdriver.Chrome(options=chrome_options)
driver.get("https://www.google.com/search?q=how+to+get+google+search+page+source+code+by+python")

try:
    # 显式等待目标元素加载,超时10秒
    search_results = WebDriverWait(driver, 10).until(
        EC.presence_of_all_elements_located((By.CLASS_NAME, "yuRUbf"))
    )
    for result in search_results:
        print(result.get_attribute('outerHTML'))
finally:
    driver.quit()

注意事项

  • 谷歌页面结构(包括class名称)会不定期更新,若yuRUbf失效,需通过浏览器开发者工具重新定位目标元素属性。
  • 频繁请求可能触发IP封禁,建议添加请求间隔,避免短时间内大量访问。

内容的提问来源于stack exchange,提问作者Saiful Islam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 02:20:37