如何为谷歌搜索Scraping代码实现翻页功能?
给谷歌搜索爬虫添加翻页功能的实现方案
谷歌搜索的翻页逻辑是通过URL中的start参数控制的:每页默认返回10条结果,第一页对应start=0,第二页start=10,第三页start=20,以此类推。基于这个规则,我们可以修改现有代码实现多页爬取。
修改后的完整代码
from urllib import response import requests import urllib import pandas as pd from requests_html import HTML from requests_html import HTMLSession def get_source(url): """Return the source code for the provided URL. Args: url (string): URL of the page to scrape. Returns: response (object): HTTP response object from requests_html. """ # 添加自定义User-Agent,避免被谷歌反爬拦截 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } try: session = HTMLSession() response = session.get(url, headers=headers) return response except requests.exceptions.RequestException as e: print(e) return None def scrape_google(query): query = urllib.parse.quote_plus(query) response = get_source("https://www.google.com/search?q=" + query) if not response: return [] links = list(response.html.absolute_links) google_domains = ('https://www.google.', 'https://google.', 'https://webcache.googleusercontent.', 'http://webcache.googleusercontent.', 'https://policies.google.', 'https://support.google.', 'https://maps.google.') for url in links[:]: if url.startswith(google_domains): links.remove(url) return links def get_results(query, start=0): query = urllib.parse.quote_plus(query) # 拼接翻页参数start url = f"https://www.google.co.uk/search?q={query}&start={start}" response = get_source(url) return response def parse_results(response): if not response: return [] css_identifier_result = ".tF2Cxc" css_identifier_title = "h3" css_identifier_link = ".yuRUbf a" css_identifier_text = ".VwiC3b" results = response.html.find(css_identifier_result) output = [] for result in results: try: item = { 'title': result.find(css_identifier_title, first=True).text, 'link': result.find(css_identifier_link, first=True).attrs['href'], 'text': result.find(css_identifier_text, first=True).text } output.append(item) except AttributeError: # 跳过解析失败的结果 continue return output def google_search(query, num_pages=1): all_results = [] for page in range(num_pages): start = page * 10 print(f"正在爬取第 {page+1} 页...") response = get_results(query, start=start) page_results = parse_results(response) all_results.extend(page_results) # 可选:添加延时,避免触发反爬 # time.sleep(2) return all_results
关键修改说明
- 添加User-Agent:在
get_source函数中加入自定义请求头,模拟浏览器访问,降低被谷歌反爬拦截的概率。 - 修改
get_results函数:新增start参数,用于拼接翻页URL。 - 增强异常处理:在
parse_results中捕获AttributeError,避免因单条结果解析失败导致整个爬虫终止;在get_source中返回None,方便上游判断请求是否成功。 - 实现多页循环:
google_search新增num_pages参数,循环生成每页的start值,请求并合并所有页面的结果。 - 可选延时:如果爬取页数较多,建议取消
time.sleep(2)的注释,给请求添加间隔,进一步降低反爬风险。
使用示例
# 爬取"python web scraping"的前3页结果 results = google_search("python web scraping", num_pages=3) # 转换为DataFrame查看 df = pd.DataFrame(results) print(df.head())
内容的提问来源于stack exchange,提问作者Julien
相关产品推荐
相关产品推荐

