如何绕过Google自动化查询安全检测批量下载Drive文件?
批量下载Google Drive PDF遭遇反爬拦截的解决方法
问题描述
有300个Google Drive的PDF文件需要下载,使用Python的requests库编写批量下载脚本,但下载30-36个文件后,Google会拦截请求并返回:
We're sorry...... but your computer or network may be sending automated queries. To protect our users, we can't process your request right now.
尝试过VPN切换IP但无效,需等待约1小时才能恢复下载,需求是绕过该安全检测。
用户原代码如下:
import requests def download_file_from_google_drive(id, destination): URL = "https://docs.google.com/uc?export=download" session = requests.Session() response = session.get(URL, params = { 'id' : id }, stream = True) if response.status_code!=200: print(response.status_code) return response.status_code print('downloading '+ destination) token = get_confirm_token(response) if token: params = { 'id' : id, 'confirm' : token } response = session.get(URL, params = params, stream = True) save_response_content(response, destination) def get_confirm_token(response): for key, value in response.cookies.items(): if key.startswith('download_warning'): return value return None def save_response_content(response, destination): CHUNK_SIZE = 32768 with open(destination, "wb") as f: i = 0 for chunk in response.iter_content(CHUNK_SIZE): print(str(i)+'%') i = i+1 if chunk: # filter out keep-alive new chunks f.write(chunk) print('downloaded '+ destination) if __name__ == "__main__": file_id = 'file id' destination = file_id+'.pdf' download_file_from_google_drive(file_id, destination)
解决方法
核心优化思路
- 模拟浏览器请求头:Google会校验请求的来源标识,添加真实浏览器的请求头可降低被识别为自动化工具的概率。
- 控制请求频率:添加随机延迟,避免短时间内大量请求触发频率阈值。
- 复用会话:使用同一个Session保持Cookie和会话状态,模拟真实用户的连续操作。
修改后的代码
import requests import time import random # 模拟主流浏览器的请求头,可根据自己的浏览器调整 HEADERS = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8', 'Accept-Language': 'zh-CN,zh;q=0.8,en-US;q=0.5,en;q=0.3', 'Referer': 'https://docs.google.com/' } # 全局复用Session,保持会话一致性 session = requests.Session() session.headers.update(HEADERS) def download_file_from_google_drive(file_id, destination): URL = "https://docs.google.com/uc?export=download" response = session.get(URL, params={'id': file_id}, stream=True) if response.status_code != 200: print(f"[{file_id}] 请求失败,状态码:{response.status_code}") return response.status_code print(f'开始下载:{destination}') token = get_confirm_token(response) if token: params = {'id': file_id, 'confirm': token} response = session.get(URL, params=params, stream=True) save_response_content(response, destination) def get_confirm_token(response): for key, value in response.cookies.items(): if key.startswith('download_warning'): return value return None def save_response_content(response, destination): CHUNK_SIZE = 32768 total_size = int(response.headers.get('content-length', 0)) downloaded = 0 with open(destination, "wb") as f: for chunk in response.iter_content(CHUNK_SIZE): if chunk: f.write(chunk) downloaded += len(chunk) # 优化进度显示,避免刷屏 if total_size > 0: progress = (downloaded / total_size) * 100 print(f'\r下载进度:{progress:.1f}%', end='') print(f'\n已完成:{destination}') # 添加2-6秒的随机延迟,模拟用户操作间隔 delay = random.uniform(2, 6) print(f'等待 {delay:.1f} 秒后继续...\n') time.sleep(delay) if __name__ == "__main__": # 替换为你的300个文件ID列表 file_ids = ['file_id_1', 'file_id_2', 'file_id_3'] for index, file_id in enumerate(file_ids): destination = f'document_{index+1}.pdf' download_file_from_google_drive(file_id, destination)
额外建议
- 如果仍然被拦截,可以尝试进一步增加延迟时间(比如5-10秒),或者分批次下载(比如每下载20个就暂停10分钟)。
- 若遇到验证码页面,可考虑使用无头浏览器模拟真实交互,但这种方式复杂度更高,建议优先调整请求频率和请求头。
内容的提问来源于stack exchange,提问作者Chandan
相关产品推荐
相关产品推荐

