You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在爬取Google搜索结果时获取完整富媒体结果?

解决Google搜索富媒体结果获取问题

你遇到的核心问题是:Google搜索的很多富媒体内容(比如社交媒体卡片、商家详细信息)是动态加载的,初始GET请求仅返回基础静态HTML,而浏览器会执行JavaScript异步拉取这些额外内容;同时Google会根据请求上下文(如Cookie、会话信息)返回不同内容,单纯用requests模拟请求缺少这些关键信息。

下面是具体解决方法:

1. 完善请求Headers与参数

Google会验证请求完整性,除User-Agent外,需补充以下Headers:

  • Cookie:从当前浏览器的Google会话中复制(注意Cookie有有效期,过期后需重新获取)
  • Accept-Language:指定语言,例如zh-CN,zh;q=0.9
  • Accept:模拟浏览器内容接受类型,例如text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8
  • Referer:设置为https://www.google.com/,模拟从主页跳转搜索

修改后的requests示例:

import requests

link = 'https://www.google.com/search?q=ellsworth+paris'
headers = {
    'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
    'Cookie': '替换为你浏览器中Google的Cookie值',
    'Accept-Language': 'zh-CN,zh;q=0.9',
    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8',
    'Referer': 'https://www.google.com/'
}

page = requests.get(link, headers=headers)
with open('out.html', 'w', encoding='utf-8') as f:
    f.write(page.text)

不过这种方法仅能解决部分静态未加载内容,完全由JS动态渲染的富媒体仍需模拟浏览器。

2. 使用浏览器自动化工具获取动态内容

用Selenium或Playwright模拟真实浏览器行为,等待JavaScript执行完成后再获取页面源码,即可拿到所有富媒体内容。

以Selenium为例,先安装依赖:

pip install selenium

下载对应浏览器驱动(如ChromeDriver)后,代码示例:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time

driver = webdriver.Chrome()  # 确保ChromeDriver路径已配置
driver.get('https://www.google.com/search?q=ellsworth+paris')

# 等待富媒体元素加载
try:
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CLASS_NAME, 'g-blk'))  # 替换为目标富媒体元素的类名
    )
except:
    pass

# 模拟滚动加载更多内容
driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
time.sleep(2)

# 获取完整页面源码
page_source = driver.page_source
with open('out_full.html', 'w', encoding='utf-8') as f:
    f.write(page_source)

driver.quit()

3. 规避反爬限制

Google有严格的反爬机制,频繁请求会触发IP封禁或验证码:

  • 每次请求后添加2-5秒的随机延时
  • 使用代理IP池分散请求来源
  • 优先使用Google官方Custom Search API(有调用限制,超出需付费,但能稳定获取结构化数据,避免反爬问题)

内容的提问来源于stack exchange,提问作者Vuk lazovic

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 22:46:18