Google Colab中requests.get遭拦截触发403错误的解决方法求助
解决Google Colab中爬取App Store RSS评论返回403的问题
本地环境运行App Store RSS评论爬取代码时所有请求均返回200,但在Google Colab中运行一段时间后被拦截,返回403错误。已尝试添加请求头、设置
time.sleep、使用代理等方法,均未解决。
可能的解决思路与优化方案
1. 实现指数退避+随机延迟的请求策略
固定时长的time.sleep容易被识别为爬虫,改用指数退避逻辑(失败次数越多,等待时间越长),同时在正常请求间加入随机小延迟,模拟人类访问节奏:
- 正常请求后添加
time.sleep(random.uniform(1, 3)) - 403失败时,等待时间设为
2 ** failed_try + random.uniform(0, 1),避免机械性重复请求
2. 完善并随机化请求头
Apple的反爬机制会校验请求头的完整性与真实性,建议:
- 构建多个真实的User-Agent池,每次请求随机选取一个
- 添加
Referer、Accept-Language、Connection等字段,模拟浏览器原生请求 - 移除GET请求中不必要的
Content-Type字段(该字段仅用于POST请求,反而可能暴露爬虫特征)
3. 优化代理使用逻辑
免费代理大多不稳定,且可能已被Apple列入黑名单:
- 改用付费代理池,确保代理IP的新鲜度与匿名性
- 每次遇到403时自动切换代理,而非一直复用同一个代理
- 先验证代理有效性再发起请求,避免无效代理浪费请求次数
4. 复用请求会话(Session)
使用requests.Session()复用TCP连接,减少请求特征,同时自动维持Cookie:
- 会话级别的请求头设置,避免重复构建
- 降低服务器端的连接检测概率
5. 规避Colab环境的IP特征
Google Colab的IP段可能被Apple标记为高风险,可尝试:
- 定期重启Colab虚拟机,更换出口IP
- 结合Tor代理(需在Colab中配置Tor环境),隐藏真实IP
6. 统一请求区域参数
你的代码中爬取评论时固定使用us区域,而前面的应用列表是按country变量切换区域,这种跨区域请求可能触发反爬,建议评论请求也使用当前country参数:
reviews_response = requesturl(f"https://itunes.apple.com/{country}/rss/customerreviews/page={n}/id={game_id}/sortBy=mostRecent/json")
修改后的示例代码
import re import glob import requests import time import random from fake_useragent import UserAgent countries = ["us", "dz", "ao", "ai", "ag"] # 代理池(示例,建议用付费代理API动态获取) proxy_pool = [ {"http": 'http://66.11.117.253:9999'}, {"http": 'http://95.154.76.20:3128'}, # 添加更多有效代理 ] failed_try = 0 # 初始化会话,复用连接 session = requests.Session() # 初始化User-Agent池 ua = UserAgent() def get_random_proxy(): return random.choice(proxy_pool) def requesturl(url): global failed_try while True: # 随机切换User-Agent和代理 headers = { 'User-Agent': ua.random, 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8', 'Referer': 'https://www.apple.com/', 'Accept-Language': 'en-US,en;q=0.5', 'Connection': 'keep-alive' } proxy = get_random_proxy() try: response = session.get(url, headers=headers, proxies=proxy, timeout=10) if response.status_code == 403: print(f"{response} Retrying request... (attempt {failed_try+1})") failed_try += 1 # 指数退避+随机延迟 sleep_time = 2 ** failed_try + random.uniform(0, 1) time.sleep(sleep_time) if failed_try >= 10: print("Too many failed attempts, skipping this URL") return None continue # 正常请求后添加随机延迟 time.sleep(random.uniform(1, 3)) failed_try = 0 return response except requests.exceptions.RequestException as e: print(f"Request error: {e}, switching proxy...") failed_try += 1 time.sleep(2) continue def download(country): genres = { "action": "7001", "strategy": "7017", "sports": "7016", } country_pos = f"{countries.index(country)+1}/{len(countries)}" print(f"\n\n\n================Scraping Country {country_pos}================") for payment in ["topfreeapplications", "toppaidapplications"]: print(f"\n\n================Scraping {payment} [{country_pos}]================") for genre in genres: print(f"\n Getting {genre} list...") genre_link = f"https://itunes.apple.com/{country}/rss/{payment}/limit=200/genre={genres[genre]}/json" dict_genre_resp = requesturl(genre_link) if not dict_genre_resp: continue print("this>", dict_genre_resp, genre_link) try: dict_genre_response = dict_genre_resp.json() except: print("Invalid JSON response, skipping...") continue for game in dict_genre_response["feed"]["entry"]: game_id = game["id"]["attributes"]["im:id"] game_name = game["title"]['label'] for n in range(1,11): # 改用当前country的评论链接 reviews_url = f"https://itunes.apple.com/{country}/rss/customerreviews/page={n}/id={game_id}/sortBy=mostRecent/json" reviews_response = requesturl(reviews_url) print(game_name, reviews_response) print() def main(): for country in countries: download(country) if __name__ == "__main__": main()
注意事项
- 需先在Colab中安装依赖库:
!pip install fake_useragent - 付费代理池建议选择支持自动切换、高匿名的服务,避免使用公开免费代理
- 爬取时遵守Apple的Robots协议与使用条款,避免过度请求对服务器造成压力
内容的提问来源于stack exchange,提问作者Aeiddius
相关产品推荐
相关产品推荐

