如何用ScrapingRobot API获取结构化JSON格式的谷歌搜索结果?
解决ScrapingRobot API获取谷歌SERP结构化JSON的问题
可行的解决方向:
尝试添加SERP解析触发参数
虽然官方文档只列出两个HTML模块,但主页明确展示了SERP结构化输出,大概率存在隐藏参数触发解析逻辑。试试在请求payload里加入"parse": "serp"或"outputType": "serp"这类字段,示例:payload = { "url": f"https://www.google.com/search?q={requests.utils.quote(SEARCH_TERM)}&hl=en&gl=us", "module": "HtmlChromeScraper", "parse": "serp" # 新增触发SERP解析的参数 }这类隐藏参数是很多爬虫API的常规操作,用来匹配主页展示的示例功能。
规范谷歌搜索URL格式
确保请求的谷歌URL是标准桌面版链接,加上语言、地区参数(比如&hl=en、&gl=us)。部分API只会对格式标准的谷歌搜索页面做结构化提取,参数混乱可能导致默认返回原始HTML。直接联系官方支持
既然主页有SERP示例但文档未更新,最直接的方式是找ScrapingRobot的客服或技术支持,询问获取SERP结构化JSON的正确模块名称或参数配置。文档滞后是这类工具的常见问题,官方能给出准确答案。临时方案:自行解析HTML
如果以上方法都无效,可直接对返回的HTML做解析提取。用BeautifulSoup实现的示例代码:from bs4 import BeautifulSoup import requests TOKEN = "..." SEARCH_TERM = "Lionel Messi" url = f"https://api.scrapingrobot.com/?token={TOKEN}" payload = { "url": f"https://www.google.com/search?q={requests.utils.quote(SEARCH_TERM)}&hl=en", "module": "HtmlChromeScraper" } headers = {"accept": "application/json", "content-type": "application/json"} response = requests.post(url, json=payload, headers=headers) if response.status_code == 200: soup = BeautifulSoup(response.json()["result"], "html.parser") serp_results = [] for result_item in soup.select("div.MjjYud"): title = result_item.select_one("h3").get_text(strip=True) if result_item.select_one("h3") else None raw_link = result_item.select_one("a")["href"] if result_item.select_one("a") else None clean_link = raw_link.replace("/url?q=", "").split("&")[0] if raw_link else None snippet = result_item.select_one("div.VwiC3b").get_text(strip=True) if result_item.select_one("div.VwiC3b") else None if title and clean_link: serp_results.append({ "title": title, "url": clean_link, "snippet": snippet }) print("提取的结构化SERP结果:\n", serp_results[:5]) else: print("Error:", response.status_code, response.text[:2000])
内容的提问来源于stack exchange,提问作者AtiehCodes
相关产品推荐
相关产品推荐

