You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将网页抓取数据流直接上传至Google Drive(无需本地存储)

问题

我有一个可将抓取的网页(支持HTML、PDF、DOCX等格式)本地保存的Python脚本,运行状态良好。此前已使用过Google Drive API的本地文件上传脚本,但希望跳过本地存储步骤,直接将requests获取的网页数据流(response.text、response.content等)上传至Google Drive,请问该如何实现?

现有脚本

本地保存网页脚本

def parse_page(self, link, q):
    response = requests.get(link, headers=self.headers, stream=True)
    print(f"\t- Response [{response.status_code}] for: {link}")
    published_date = self.get_date(response.text)
    if "text/html" in response.headers["content-type"]:
        format = ".html"
        data = response.text
        mode = "w"
    elif "application/pdf" in response.headers["content-type"]:
        format = ".pdf"
        data = response.content
        mode = "wb"
    elif "officedocument.word" in response.headers["content-type"]:
        format = ".docx"
        data = response
        mode = "wb"
    folder = urlparse(link).netloc
    filename = urlparse(link).path.split("/")[-1].replace("\n","")
    if filename == '':
        filename = folder + datetime.now().strftime("_%H_%M_%S")
    # Windows路径
    path = self.root_folder + "\\" + folder + "\\" + filename + format
    with self.safe_open_w(path, mode) as f:
        if format == '.docx':
            for chunk in data.iter_content(16*1024):  # 16KB chunks
                f.write(chunk)
        else:
            f.write(data)

Google Drive本地文件上传脚本

try:
    # 调用Drive v3 API
    file_metadata = {'name': 'file.jpg'}
    media = MediaFileUpload('file.jpg', mimetype='image/jpeg')
    # pylint: disable=maybe-no-member
    file = service.files().create(body=file_metadata, media_body=media, fields='id').execute()
    print(f'File ID: {file.get("id")}')

except HttpError as error:
    print(f'An error occurred: {error}')
    file = None
except Exception as e:
    print(e)
实现方案

Google Drive API提供的MediaIoBaseUpload类支持直接上传内存中的数据流,无需先写入本地文件。我们可以基于现有抓取逻辑,结合该类实现流式上传:

步骤1:导入必要模块

除原有的依赖外,补充导入内存流处理和Drive上传相关模块:

import io
from googleapiclient.http import MediaIoBaseUpload

步骤2:改造数据流处理逻辑

将原本地保存的代码替换为Drive上传逻辑,针对不同内容类型处理数据流:

  • HTML文本:将response.text转为字节流,包装成io.BytesIO对象
  • PDF文件:直接用response.content包装成io.BytesIO
  • DOCX文件:利用response.raw获取原始字节流,需重置流指针到起始位置

完整整合代码

以下是改造后的parse_page方法,整合了Drive上传逻辑(假设self.service是已初始化的Google Drive API服务对象):

def parse_page(self, link, q):
    response = requests.get(link, headers=self.headers, stream=True)
    print(f"\t- Response [{response.status_code}] for: {link}")
    published_date = self.get_date(response.text)
    
    # 确定文件类型、MIME类型和数据流
    file_format = ""
    mime_type = ""
    stream = None
    
    if "text/html" in response.headers["content-type"]:
        file_format = ".html"
        mime_type = "text/html"
        # 将文本转为字节流
        stream = io.BytesIO(response.text.encode('utf-8'))
    elif "application/pdf" in response.headers["content-type"]:
        file_format = ".pdf"
        mime_type = "application/pdf"
        stream = io.BytesIO(response.content)
    elif "officedocument.word" in response.headers["content-type"]:
        file_format = ".docx"
        mime_type = "application/vnd.openxmlformats-officedocument.wordprocessingml.document"
        # 使用原始响应流,重置指针到开头
        stream = response.raw
        stream.seek(0)
    
    # 生成文件名
    folder = urlparse(link).netloc
    filename = urlparse(link).path.split("/")[-1].replace("\n","")
    if filename == '':
        filename = folder + datetime.now().strftime("_%H_%M_%S")
    full_filename = filename + file_format
    
    # 上传到Google Drive
    try:
        file_metadata = {'name': full_filename}
        # 创建Media对象,启用分块上传(支持大文件断点续传)
        media = MediaIoBaseUpload(stream, mimetype=mime_type, chunksize=16*1024, resumable=True)
        # 调用API完成上传
        file = self.service.files().create(body=file_metadata, media_body=media, fields='id').execute()
        print(f"\t- Uploaded to Drive, File ID: {file.get('id')}")
    except HttpError as error:
        print(f"\t- Drive upload error: {error}")
    except Exception as e:
        print(f"\t- Upload failed: {e}")

关键说明

  1. 分块上传:设置chunksize=16*1024和resumable=True,既减少内存占用,又支持大文件断点续传
  2. 流指针重置:使用response.raw时必须调用stream.seek(0),确保从流的起始位置读取数据
  3. MIME类型匹配:每个文件类型的MIME类型需与Google Drive兼容,避免文件识别错误

内容的提问来源于stack exchange,提问作者Mauro

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 03:54:52