如何解析谷歌搜索结果统计数字?爬取异常排查与解决
谷歌搜索结果统计数字解析问题及解决办法
问题现象
尝试用Python解析谷歌搜索结果中的统计数字,编写了如下代码:
import requests from bs4 import BeautifulSoup # 发起谷歌搜索请求 response = requests.get("https://www.google.com/search?q=book") soup = BeautifulSoup(response.content, 'html.parser') phrase_extract = soup.find_all(id="result-stats") print(phrase_extract)
执行后输出为空列表:
[]
在Chrome浏览器查看网页源码时,能明确看到目标元素:
<div id="result-stats">About 13,120,000,000 results<nobr> (0.51 seconds) </nobr></div>
但打印response.text时却找不到该字符串。
问题原因
谷歌会对请求的User-Agent进行检测,requests库默认的请求头标识会被判定为非浏览器请求,返回的页面内容与浏览器实际加载的内容不一致,因此无法定位到目标元素。
解决办法
在请求中添加模拟浏览器的User-Agent请求头,让谷歌返回与浏览器一致的页面内容。更新后的代码如下:
import requests from bs4 import BeautifulSoup # 发起谷歌搜索请求 response = requests.get("https://www.google.com/search?q=book", headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36 Edg/91.0.864.59'}) soup = BeautifulSoup(response.content, 'html.parser') phrase_extract = soup.find_all(id="result-stats") print(phrase_extract)
添加headers参数后即可正常获取目标元素。
内容的提问来源于stack exchange,提问作者Mark Kang
相关产品推荐
相关产品推荐

