Python谷歌搜索爬虫添加分页功能实现多页结果抓取求助
谷歌搜索爬虫分页功能实现方案
核心实现逻辑:谷歌搜索的分页通过URL的start参数控制,每页默认返回10条搜索结果,第n页对应的start参数值为(n-1)*10,只需要在原有代码基础上新增分页参数遍历、多页结果合并逻辑即可,完全保留原有的标题、链接、描述抓取规则,不会出现仅返回URL的问题。
改动说明
- 为
get_results函数新增start参数,用于构造分页搜索URL - 为
google_search函数新增page_num参数,支持自定义爬取总页数,默认值为1兼容原有单页逻辑 - 新增循环遍历逻辑,逐页抓取并合并所有结果
- 优化解析逻辑的鲁棒性,避免部分字段缺失时触发报错
完整修改后代码
import requests import urllib import pandas as pd import time import random from requests_html import HTML from requests_html import HTMLSession def get_source(url): """Return the source code for the provided URL. Args: url (string): URL of the page to scrape. Returns: response (object): HTTP response object from requests_html. """ try: session = HTMLSession() # 添加请求头模拟浏览器,降低被谷歌拦截的概率 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" } response = session.get(url, headers=headers) return response except requests.exceptions.RequestException as e: print(e) return None def get_results(query, start=0): query = urllib.parse.quote_plus(query) # 拼接start参数实现分页 search_url = f"https://www.google.co.uk/search?q={query}&start={start}" response = get_source(search_url) return response def parse_results(response): if not response: return [] css_identifier_result = ".tF2Cxc" css_identifier_title = "h3" css_identifier_link = ".yuRUbf a" css_identifier_text = ".IsZvec" results = response.html.find(css_identifier_result) output = [] for result in results: item = { 'title': result.find(css_identifier_title, first=True).text if result.find(css_identifier_title, first=True) else '', 'link': result.find(css_identifier_link, first=True).attrs['href'] if result.find(css_identifier_link, first=True) else '', 'text': result.find(css_identifier_text, first=True).text if result.find(css_identifier_text, first=True) else '' } output.append(item) return output def google_search(query, page_num=3): all_results = [] for page in range(page_num): start = page * 10 response = get_results(query, start) page_res = parse_results(response) if not page_res: # 无结果时提前终止爬取 break all_results.extend(page_res) # 翻页添加随机延迟,避免触发反爬 time.sleep(random.uniform(1,3)) return all_results # 使用示例 query = input("Enter your value: ") results = google_search(query, page_num=3) # 打印结果查看 for res in results: print(res)
注意事项
- 如果爬取页数较多,可根据需求调整
google_search函数的延迟时长,进一步降低被拦截概率 - 谷歌搜索的CSS选择器可能会随页面更新调整,如果出现无结果返回的情况,可检查对应选择器是否匹配最新页面结构
内容的提问来源于stack exchange,提问作者Raspberry Lemon
相关产品推荐
相关产品推荐

