You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python谷歌搜索爬虫添加分页功能实现多页结果抓取求助

谷歌搜索爬虫分页功能实现方案

核心实现逻辑:谷歌搜索的分页通过URL的start参数控制,每页默认返回10条搜索结果,第n页对应的start参数值为(n-1)*10,只需要在原有代码基础上新增分页参数遍历、多页结果合并逻辑即可,完全保留原有的标题、链接、描述抓取规则,不会出现仅返回URL的问题。

改动说明

  • 为get_results函数新增start参数,用于构造分页搜索URL
  • 为google_search函数新增page_num参数,支持自定义爬取总页数,默认值为1兼容原有单页逻辑
  • 新增循环遍历逻辑,逐页抓取并合并所有结果
  • 优化解析逻辑的鲁棒性,避免部分字段缺失时触发报错

完整修改后代码

import requests
import urllib
import pandas as pd
import time
import random
from requests_html import HTML
from requests_html import HTMLSession


def get_source(url):
    """Return the source code for the provided URL. 

    Args: 
        url (string): URL of the page to scrape.

    Returns:
        response (object): HTTP response object from requests_html. 
    """
    try:
        session = HTMLSession()
        # 添加请求头模拟浏览器,降低被谷歌拦截的概率
        headers = {
            "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36"
        }
        response = session.get(url, headers=headers)
        return response
    except requests.exceptions.RequestException as e:
        print(e)
        return None

def get_results(query, start=0):
    query = urllib.parse.quote_plus(query)
    # 拼接start参数实现分页
    search_url = f"https://www.google.co.uk/search?q={query}&start={start}"
    response = get_source(search_url)
    return response

def parse_results(response):
    if not response:
        return []
    css_identifier_result = ".tF2Cxc"
    css_identifier_title = "h3"
    css_identifier_link = ".yuRUbf a"
    css_identifier_text = ".IsZvec"
    
    results = response.html.find(css_identifier_result)
    output = []
    
    for result in results:
        item = {
            'title': result.find(css_identifier_title, first=True).text if result.find(css_identifier_title, first=True) else '',
            'link': result.find(css_identifier_link, first=True).attrs['href'] if result.find(css_identifier_link, first=True) else '',
            'text': result.find(css_identifier_text, first=True).text if result.find(css_identifier_text, first=True) else ''
        }
        output.append(item)
    return output

def google_search(query, page_num=3):
    all_results = []
    for page in range(page_num):
        start = page * 10
        response = get_results(query, start)
        page_res = parse_results(response)
        if not page_res:
            # 无结果时提前终止爬取
            break
        all_results.extend(page_res)
        # 翻页添加随机延迟,避免触发反爬
        time.sleep(random.uniform(1,3))
    return all_results

# 使用示例
query = input("Enter your value: ")
results = google_search(query, page_num=3)
# 打印结果查看
for res in results:
    print(res)

注意事项

  • 如果爬取页数较多,可根据需求调整google_search函数的延迟时长,进一步降低被拦截概率
  • 谷歌搜索的CSS选择器可能会随页面更新调整,如果出现无结果返回的情况,可检查对应选择器是否匹配最新页面结构

内容的提问来源于stack exchange,提问作者Raspberry Lemon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 06:24:05