You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Google Drive API Python按扩展名筛选下载文件报错排查

Google Drive指定文件下载方案修复

问题背景

  • 需求:从两个独立Google Drive文件夹分别下载目标文件:PDF文件以Quote为文件名前缀、后缀为.pdf(例:Quote123456.pdf),XML文件以FGT为文件名前缀、后缀为.xml(例:FGT1236098.xml)
  • 原有代码在仅存放目标文件的测试文件夹可正常运行,但生产文件夹存在其他无关文件、子文件夹,需要实现精准过滤仅下载目标文件
  • 原有代码运行先后触发两类报错:
    • 第一类:googleapiclient.errors.HttpError 400,提示fields参数存在无效字段选择
    • 第二类:调整代码后触发KeyError: 'createdTime'

问题根因

  1. 400报错核心原因:过滤逻辑放错参数位置。name contains 'Quote'这类文件筛选规则属于q(查询参数)的接收内容,而fields参数仅用于声明接口需要返回的文件属性字段,不能写入过滤逻辑。
  2. 分页逻辑存在两处错误:一是方法名拼写错误files.extende应为files.extend,导致分页拉取的文件无法追加到结果列表;二是分页请求未传入pageToken参数,会导致重复拉取第一页内容,无法获取全量文件。
  3. KeyError报错核心原因:一是过滤逻辑错误导致接口未正常返回createdTime字段;二是未过滤子文件夹、无关文件,当返回结果为空时生成的DataFrame不存在对应列,直接取列就会触发键错误。
  4. 额外遗漏:原有查询未排除回收站文件、未过滤子文件夹,会拉取到不需要的内容。
  5. 语法说明:Google Drive API的name contains查询不需要手动添加通配符,也不需要手动做URL编码,官方Python SDK会自动处理请求编码。

修正后完整代码

import google_drive.constants as c
import os
import io
import pandas as pd
from googleapiclient.http import MediaIoBaseDownload

service = Create_Service(c.CLIENT_SECRET_FILE, c.API_NAME, c.API_VERSION, c.SCOPES)


def compare_file_dates(df):
    # 转换时间格式
    df['createdTime'] = pd.to_datetime(df['createdTime'])
    # 取最新创建的文件
    row = df[df.createdTime == df.createdTime.max()]
    latest = row['id'].iloc[0]
    print(f"找到最新目标文件ID:{latest}")
    return latest


def get_latest_file(folder_id, name_prefix, file_suffix):
    # 过滤条件写在q参数:指定父文件夹、排除回收站、排除文件夹、匹配前缀和后缀
    query = f"""
    parents = '{folder_id}' 
    and trashed = false 
    and mimeType != 'application/vnd.google-apps.folder' 
    and name contains '{name_prefix}' 
    and name contains '.{file_suffix}'
    """
    # fields仅声明需要返回的字段,不写过滤逻辑
    fields = "nextPageToken, files(id, name, createdTime, mimeType)"
    response = service.files().list(q=query, fields=fields).execute()
    files = response.get('files', [])
    next_page_token = response.get('nextPageToken')

    # 修复分页逻辑
    while next_page_token:
        response = service.files().list(
            q=query, 
            fields=fields, 
            pageToken=next_page_token
        ).execute()
        files.extend(response.get('files', []))
        next_page_token = response.get('nextPageToken')

    # 空结果判断
    if not files:
        raise ValueError(f"文件夹{folder_id}下未找到匹配{name_prefix}前缀、{file_suffix}后缀的文件")
    
    df = pd.DataFrame(files)
    latest_file = compare_file_dates(df)
    return latest_file


def download_files():
    # 传入前缀、后缀参数,分别获取两个文件夹下的最新目标文件
    pdf_id = get_latest_file(c.PDF_FOLDER_ID, name_prefix='Quote', file_suffix='pdf')
    xml_id = get_latest_file(c.XML_FOLDER_ID, name_prefix='FGT', file_suffix='xml')
    file_ids = [pdf_id, xml_id]
    file_names = ['expense.pdf', 'sell.xml']

    # 确保下载目录存在
    os.makedirs('google_drive/downloads', exist_ok=True)

    for file_id, file_name in zip(file_ids, file_names):
        request = service.files().get_media(fileId=file_id)
        fh = io.BytesIO()
        downloader = MediaIoBaseDownload(fd=fh, request=request)
        done = False

        while not done:
            status, done = downloader.next_chunk()
            print(f'{file_name} 下载进度:{status.progress() * 100:.2f}%')

        fh.seek(0)
        save_path = os.path.join('google_drive/downloads', file_name)
        with open(save_path, 'wb') as f:
            f.write(fh.read())
        print(f'{file_name} 已保存到:{save_path}')


if __name__ == '__main__':
    download_files()

关键修正说明

  • 所有过滤逻辑统一放在q参数中,新增trashed = false排除回收站文件、mimeType != 'application/vnd.google-apps.folder'过滤子文件夹,确保仅返回符合命名规则的目标格式文件
  • fields参数仅保留需要返回的属性字段,不再写入过滤逻辑,解决400参数错误
  • 修复分页逻辑的拼写错误、缺失pageToken的问题,确保能拉取到文件夹下全量符合条件的文件
  • 新增空结果判断,未找到匹配文件时抛出明确提示,避免无意义的KeyError
  • 新增下载目录自动创建逻辑,避免目录不存在导致的写入错误
  • 不需要手动添加通配符、不需要手动处理URL编码,SDK会自动完成相关处理

内容的提问来源于stack exchange,提问作者Nick

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.30 02:39:16