You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为谷歌搜索Scraping代码实现翻页功能?

给谷歌搜索爬虫添加翻页功能的实现方案

谷歌搜索的翻页逻辑是通过URL中的start参数控制的:每页默认返回10条结果,第一页对应start=0,第二页start=10,第三页start=20,以此类推。基于这个规则,我们可以修改现有代码实现多页爬取。

修改后的完整代码

from urllib import response
import requests
import urllib
import pandas as pd
from requests_html import HTML
from requests_html import HTMLSession

def get_source(url):
    """Return the source code for the provided URL. 

    Args: 
        url (string): URL of the page to scrape.

    Returns:
        response (object): HTTP response object from requests_html. 
    """
    # 添加自定义User-Agent,避免被谷歌反爬拦截
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
    }
    try:
        session = HTMLSession()
        response = session.get(url, headers=headers)
        return response

    except requests.exceptions.RequestException as e:
        print(e)
        return None
        
def scrape_google(query):
    query = urllib.parse.quote_plus(query)
    response = get_source("https://www.google.com/search?q=" + query)

    if not response:
        return []
        
    links = list(response.html.absolute_links)
    google_domains = ('https://www.google.', 
                      'https://google.', 
                      'https://webcache.googleusercontent.', 
                      'http://webcache.googleusercontent.', 
                      'https://policies.google.',
                      'https://support.google.',
                      'https://maps.google.')

    for url in links[:]:
        if url.startswith(google_domains):
            links.remove(url)

    return links

def get_results(query, start=0):
    query = urllib.parse.quote_plus(query)
    # 拼接翻页参数start
    url = f"https://www.google.co.uk/search?q={query}&start={start}"
    response = get_source(url)
    return response

def parse_results(response):
    if not response:
        return []
        
    css_identifier_result = ".tF2Cxc"
    css_identifier_title = "h3"
    css_identifier_link = ".yuRUbf a"
    css_identifier_text = ".VwiC3b"
    
    results = response.html.find(css_identifier_result)

    output = []
    
    for result in results:
        try:
            item = {
                'title': result.find(css_identifier_title, first=True).text,
                'link': result.find(css_identifier_link, first=True).attrs['href'],
                'text': result.find(css_identifier_text, first=True).text
            }
            output.append(item)
        except AttributeError:
            # 跳过解析失败的结果
            continue
        
    return output

def google_search(query, num_pages=1):
    all_results = []
    for page in range(num_pages):
        start = page * 10
        print(f"正在爬取第 {page+1} 页...")
        response = get_results(query, start=start)
        page_results = parse_results(response)
        all_results.extend(page_results)
        # 可选:添加延时,避免触发反爬
        # time.sleep(2)
    return all_results

关键修改说明

  1. 添加User-Agent:在get_source函数中加入自定义请求头,模拟浏览器访问,降低被谷歌反爬拦截的概率。
  2. 修改get_results函数:新增start参数,用于拼接翻页URL。
  3. 增强异常处理:在parse_results中捕获AttributeError,避免因单条结果解析失败导致整个爬虫终止;在get_source中返回None,方便上游判断请求是否成功。
  4. 实现多页循环:google_search新增num_pages参数,循环生成每页的start值,请求并合并所有页面的结果。
  5. 可选延时:如果爬取页数较多,建议取消time.sleep(2)的注释,给请求添加间隔,进一步降低反爬风险。

使用示例

# 爬取"python web scraping"的前3页结果
results = google_search("python web scraping", num_pages=3)
# 转换为DataFrame查看
df = pd.DataFrame(results)
print(df.head())

内容的提问来源于stack exchange,提问作者Julien

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 10:24:13