Google Drive API Python按扩展名筛选下载文件报错排查
Google Drive指定文件下载方案修复
问题背景
- 需求:从两个独立Google Drive文件夹分别下载目标文件:PDF文件以
Quote为文件名前缀、后缀为.pdf(例:Quote123456.pdf),XML文件以FGT为文件名前缀、后缀为.xml(例:FGT1236098.xml) - 原有代码在仅存放目标文件的测试文件夹可正常运行,但生产文件夹存在其他无关文件、子文件夹,需要实现精准过滤仅下载目标文件
- 原有代码运行先后触发两类报错:
- 第一类:
googleapiclient.errors.HttpError 400,提示fields参数存在无效字段选择 - 第二类:调整代码后触发
KeyError: 'createdTime'
- 第一类:
问题根因
- 400报错核心原因:过滤逻辑放错参数位置。
name contains 'Quote'这类文件筛选规则属于q(查询参数)的接收内容,而fields参数仅用于声明接口需要返回的文件属性字段,不能写入过滤逻辑。 - 分页逻辑存在两处错误:一是方法名拼写错误
files.extende应为files.extend,导致分页拉取的文件无法追加到结果列表;二是分页请求未传入pageToken参数,会导致重复拉取第一页内容,无法获取全量文件。 - KeyError报错核心原因:一是过滤逻辑错误导致接口未正常返回
createdTime字段;二是未过滤子文件夹、无关文件,当返回结果为空时生成的DataFrame不存在对应列,直接取列就会触发键错误。 - 额外遗漏:原有查询未排除回收站文件、未过滤子文件夹,会拉取到不需要的内容。
- 语法说明:Google Drive API的
name contains查询不需要手动添加通配符,也不需要手动做URL编码,官方Python SDK会自动处理请求编码。
修正后完整代码
import google_drive.constants as c import os import io import pandas as pd from googleapiclient.http import MediaIoBaseDownload service = Create_Service(c.CLIENT_SECRET_FILE, c.API_NAME, c.API_VERSION, c.SCOPES) def compare_file_dates(df): # 转换时间格式 df['createdTime'] = pd.to_datetime(df['createdTime']) # 取最新创建的文件 row = df[df.createdTime == df.createdTime.max()] latest = row['id'].iloc[0] print(f"找到最新目标文件ID:{latest}") return latest def get_latest_file(folder_id, name_prefix, file_suffix): # 过滤条件写在q参数:指定父文件夹、排除回收站、排除文件夹、匹配前缀和后缀 query = f""" parents = '{folder_id}' and trashed = false and mimeType != 'application/vnd.google-apps.folder' and name contains '{name_prefix}' and name contains '.{file_suffix}' """ # fields仅声明需要返回的字段,不写过滤逻辑 fields = "nextPageToken, files(id, name, createdTime, mimeType)" response = service.files().list(q=query, fields=fields).execute() files = response.get('files', []) next_page_token = response.get('nextPageToken') # 修复分页逻辑 while next_page_token: response = service.files().list( q=query, fields=fields, pageToken=next_page_token ).execute() files.extend(response.get('files', [])) next_page_token = response.get('nextPageToken') # 空结果判断 if not files: raise ValueError(f"文件夹{folder_id}下未找到匹配{name_prefix}前缀、{file_suffix}后缀的文件") df = pd.DataFrame(files) latest_file = compare_file_dates(df) return latest_file def download_files(): # 传入前缀、后缀参数,分别获取两个文件夹下的最新目标文件 pdf_id = get_latest_file(c.PDF_FOLDER_ID, name_prefix='Quote', file_suffix='pdf') xml_id = get_latest_file(c.XML_FOLDER_ID, name_prefix='FGT', file_suffix='xml') file_ids = [pdf_id, xml_id] file_names = ['expense.pdf', 'sell.xml'] # 确保下载目录存在 os.makedirs('google_drive/downloads', exist_ok=True) for file_id, file_name in zip(file_ids, file_names): request = service.files().get_media(fileId=file_id) fh = io.BytesIO() downloader = MediaIoBaseDownload(fd=fh, request=request) done = False while not done: status, done = downloader.next_chunk() print(f'{file_name} 下载进度:{status.progress() * 100:.2f}%') fh.seek(0) save_path = os.path.join('google_drive/downloads', file_name) with open(save_path, 'wb') as f: f.write(fh.read()) print(f'{file_name} 已保存到:{save_path}') if __name__ == '__main__': download_files()
关键修正说明
- 所有过滤逻辑统一放在
q参数中,新增trashed = false排除回收站文件、mimeType != 'application/vnd.google-apps.folder'过滤子文件夹,确保仅返回符合命名规则的目标格式文件 fields参数仅保留需要返回的属性字段,不再写入过滤逻辑,解决400参数错误- 修复分页逻辑的拼写错误、缺失
pageToken的问题,确保能拉取到文件夹下全量符合条件的文件 - 新增空结果判断,未找到匹配文件时抛出明确提示,避免无意义的KeyError
- 新增下载目录自动创建逻辑,避免目录不存在导致的写入错误
- 不需要手动添加通配符、不需要手动处理URL编码,SDK会自动完成相关处理
内容的提问来源于stack exchange,提问作者Nick
相关产品推荐
相关产品推荐

