You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何借助代理在Python中高效获取谷歌搜索结果URL列表?

如何高效借助代理获取谷歌搜索结果URL列表?

问题背景

原本使用googlesearch库快速获取搜索结果:

from googlesearch import search
list(search(f"{query}", num_results))

但频繁触发429请求过多错误:

requests.exceptions.HTTPError: 429 Client Error: Too Many Requests for url: https://www.google.com/sorry/index?continue=https://www.google.com/search%3Fq%{query}%26num%3D10%26hl%3Den%26start%3D67&hl=en&q=EhAmABcACfAIIME0fDvEUYF8GOKX1KQGIjAEGg2nloeEEAcko9umYCP9uPHRWoSo2odE3n3ZgbQ1L6lDvGfyai6798pyy3iU5vcyAXJaAUM

自行用requests+BeautifulSoup实现的方案效率极低,获取100条URL耗时约1小时,远不如原方法的1秒速度:

search_results = []
retry = True
while retry:
    try:
        response = requests.get(f"https://www.google.com/search?q={query}", 
                                    headers={
                                        'User-Agent': user_agent,
                                        'Referer': 'https://www.google.com/',
                                        'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.9', 
                                        'Accept-Encoding': 'gzip, deflate, br', 
                                        'Accept-Language': 'en-US,en;q=0.9,en-gb', 
                                    }, 
                                    proxies={
                                        "http": proxy, 
                                        "https": proxy}, 
                                    timeout=TIMEOUT_THRESHOLD*2)
    
        if response.status_code == 200:
            soup = BeautifulSoup(response.content, 'html.parser')
            
            for link in soup.select('.yuRUbf a'):
                url = link['href']
                
                search_results.append(url)
                
                if len(search_results) >= num_results:
                    retry = False
                    break  
        
        else:
            proxy = get_working_proxy(proxies)
            user_agent = random.choice(user_agents)

    except Exception as e:
        proxy = get_working_proxy(proxies)
        user_agent = random.choice(user_agents)
        print(f"An error occurred in tips search: {str(e)}")   

优化方案

1. 给googlesearch库配置代理与请求参数

原库本身支持通过参数配置代理、User-Agent和请求间隔,直接调整就能避免429错误,同时保留原有的高效性:

from googlesearch import search
import random

# 配置代理、随机UA和请求间隔
search_results = list(
    search(
        query,
        num_results=num_results,
        user_agent=random.choice(user_agents),  # 传入随机UA列表中的一个
        proxy=proxy,  # 传入有效的代理地址
        pause=2,  # 请求间隔,避免触发反爬,可根据情况调整
        num=10  # 每页结果数,默认10,和谷歌页面一致
    )
)

如果单个代理不够用,可以实现代理轮换逻辑:每次请求失败时切换代理,重新调用search方法。

2. 用并发请求优化自定义爬虫

你的自定义方案效率低的核心原因是单线程串行请求,改用concurrent.futures.ThreadPoolExecutor实现多线程并发,同时配合代理池和UA轮换,能大幅提升速度:

import requests
from bs4 import BeautifulSoup
import random
from concurrent.futures import ThreadPoolExecutor, as_completed

def fetch_page(page_num, query, user_agents, proxies):
    """单页搜索结果抓取"""
    start = page_num * 10
    url = f"https://www.google.com/search?q={query}&start={start}"
    headers = {
        'User-Agent': random.choice(user_agents),
        'Referer': 'https://www.google.com/',
        'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.9',
        'Accept-Encoding': 'gzip, deflate, br',
        'Accept-Language': 'en-US,en;q=0.9,en-gb',
    }
    proxy = random.choice(proxies)
    try:
        response = requests.get(url, headers=headers, proxies={"http": proxy, "https": proxy}, timeout=10)
        response.raise_for_status()
        soup = BeautifulSoup(response.content, 'html.parser')
        return [link['href'] for link in soup.select('.yuRUbf a')]
    except Exception as e:
        print(f"Page {page_num} fetch failed: {str(e)}")
        return []

def get_search_results(query, num_results, user_agents, proxies):
    """批量获取搜索结果"""
    total_pages = (num_results + 9) // 10  # 计算需要的页数
    search_results = []
    with ThreadPoolExecutor(max_workers=5) as executor:  # 线程数根据代理质量调整,避免被封
        futures = [executor.submit(fetch_page, page, query, user_agents, proxies) for page in range(total_pages)]
        for future in as_completed(futures):
            results = future.result()
            search_results.extend(results)
            if len(search_results) >= num_results:
                break
    return search_results[:num_results]

# 使用示例
# search_results = get_search_results("your query", 100, your_user_agents_list, your_proxies_list)

注意:线程数不要设置过高,否则容易触发谷歌的反爬机制;同时要确保代理池中的代理是有效且高质量的,无效代理会拖慢整体速度。

3. 使用专门的反爬友好搜索库

有些第三方库针对谷歌搜索的反爬机制做了优化,支持代理轮换、UA随机等功能,这类库能平衡速度和反爬规避,比如pygooglesearch,无需额外API密钥即可使用。

关键注意事项

  • 代理质量是核心:使用高匿代理,避免使用公开免费代理(大多无效或被谷歌拉黑),可以考虑付费代理池服务。
  • 请求频率控制:即使使用代理,也不要无间隔请求,设置合理的pause或线程休眠时间。
  • UA轮换:每次请求使用不同的User-Agent,模拟真实浏览器请求。

内容的提问来源于stack exchange,提问作者Ricardo Francois

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 08:39:59