如何借助代理在Python中高效获取谷歌搜索结果URL列表?
如何高效借助代理获取谷歌搜索结果URL列表?
问题背景
原本使用googlesearch库快速获取搜索结果:
from googlesearch import search list(search(f"{query}", num_results))
但频繁触发429请求过多错误:
requests.exceptions.HTTPError: 429 Client Error: Too Many Requests for url: https://www.google.com/sorry/index?continue=https://www.google.com/search%3Fq%{query}%26num%3D10%26hl%3Den%26start%3D67&hl=en&q=EhAmABcACfAIIME0fDvEUYF8GOKX1KQGIjAEGg2nloeEEAcko9umYCP9uPHRWoSo2odE3n3ZgbQ1L6lDvGfyai6798pyy3iU5vcyAXJaAUM
自行用requests+BeautifulSoup实现的方案效率极低,获取100条URL耗时约1小时,远不如原方法的1秒速度:
search_results = [] retry = True while retry: try: response = requests.get(f"https://www.google.com/search?q={query}", headers={ 'User-Agent': user_agent, 'Referer': 'https://www.google.com/', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.9', 'Accept-Encoding': 'gzip, deflate, br', 'Accept-Language': 'en-US,en;q=0.9,en-gb', }, proxies={ "http": proxy, "https": proxy}, timeout=TIMEOUT_THRESHOLD*2) if response.status_code == 200: soup = BeautifulSoup(response.content, 'html.parser') for link in soup.select('.yuRUbf a'): url = link['href'] search_results.append(url) if len(search_results) >= num_results: retry = False break else: proxy = get_working_proxy(proxies) user_agent = random.choice(user_agents) except Exception as e: proxy = get_working_proxy(proxies) user_agent = random.choice(user_agents) print(f"An error occurred in tips search: {str(e)}")
优化方案
1. 给googlesearch库配置代理与请求参数
原库本身支持通过参数配置代理、User-Agent和请求间隔,直接调整就能避免429错误,同时保留原有的高效性:
from googlesearch import search import random # 配置代理、随机UA和请求间隔 search_results = list( search( query, num_results=num_results, user_agent=random.choice(user_agents), # 传入随机UA列表中的一个 proxy=proxy, # 传入有效的代理地址 pause=2, # 请求间隔,避免触发反爬,可根据情况调整 num=10 # 每页结果数,默认10,和谷歌页面一致 ) )
如果单个代理不够用,可以实现代理轮换逻辑:每次请求失败时切换代理,重新调用search方法。
2. 用并发请求优化自定义爬虫
你的自定义方案效率低的核心原因是单线程串行请求,改用concurrent.futures.ThreadPoolExecutor实现多线程并发,同时配合代理池和UA轮换,能大幅提升速度:
import requests from bs4 import BeautifulSoup import random from concurrent.futures import ThreadPoolExecutor, as_completed def fetch_page(page_num, query, user_agents, proxies): """单页搜索结果抓取""" start = page_num * 10 url = f"https://www.google.com/search?q={query}&start={start}" headers = { 'User-Agent': random.choice(user_agents), 'Referer': 'https://www.google.com/', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.9', 'Accept-Encoding': 'gzip, deflate, br', 'Accept-Language': 'en-US,en;q=0.9,en-gb', } proxy = random.choice(proxies) try: response = requests.get(url, headers=headers, proxies={"http": proxy, "https": proxy}, timeout=10) response.raise_for_status() soup = BeautifulSoup(response.content, 'html.parser') return [link['href'] for link in soup.select('.yuRUbf a')] except Exception as e: print(f"Page {page_num} fetch failed: {str(e)}") return [] def get_search_results(query, num_results, user_agents, proxies): """批量获取搜索结果""" total_pages = (num_results + 9) // 10 # 计算需要的页数 search_results = [] with ThreadPoolExecutor(max_workers=5) as executor: # 线程数根据代理质量调整,避免被封 futures = [executor.submit(fetch_page, page, query, user_agents, proxies) for page in range(total_pages)] for future in as_completed(futures): results = future.result() search_results.extend(results) if len(search_results) >= num_results: break return search_results[:num_results] # 使用示例 # search_results = get_search_results("your query", 100, your_user_agents_list, your_proxies_list)
注意:线程数不要设置过高,否则容易触发谷歌的反爬机制;同时要确保代理池中的代理是有效且高质量的,无效代理会拖慢整体速度。
3. 使用专门的反爬友好搜索库
有些第三方库针对谷歌搜索的反爬机制做了优化,支持代理轮换、UA随机等功能,这类库能平衡速度和反爬规避,比如pygooglesearch,无需额外API密钥即可使用。
关键注意事项
- 代理质量是核心:使用高匿代理,避免使用公开免费代理(大多无效或被谷歌拉黑),可以考虑付费代理池服务。
- 请求频率控制:即使使用代理,也不要无间隔请求,设置合理的
pause或线程休眠时间。 - UA轮换:每次请求使用不同的User-Agent,模拟真实浏览器请求。
内容的提问来源于stack exchange,提问作者Ricardo Francois
相关产品推荐
相关产品推荐

