如何将网页抓取数据流直接上传至Google Drive(无需本地存储)
问题
我有一个可将抓取的网页(支持HTML、PDF、DOCX等格式)本地保存的Python脚本,运行状态良好。此前已使用过Google Drive API的本地文件上传脚本,但希望跳过本地存储步骤,直接将requests获取的网页数据流(response.text、response.content等)上传至Google Drive,请问该如何实现?
现有脚本
本地保存网页脚本
def parse_page(self, link, q): response = requests.get(link, headers=self.headers, stream=True) print(f"\t- Response [{response.status_code}] for: {link}") published_date = self.get_date(response.text) if "text/html" in response.headers["content-type"]: format = ".html" data = response.text mode = "w" elif "application/pdf" in response.headers["content-type"]: format = ".pdf" data = response.content mode = "wb" elif "officedocument.word" in response.headers["content-type"]: format = ".docx" data = response mode = "wb" folder = urlparse(link).netloc filename = urlparse(link).path.split("/")[-1].replace("\n","") if filename == '': filename = folder + datetime.now().strftime("_%H_%M_%S") # Windows路径 path = self.root_folder + "\\" + folder + "\\" + filename + format with self.safe_open_w(path, mode) as f: if format == '.docx': for chunk in data.iter_content(16*1024): # 16KB chunks f.write(chunk) else: f.write(data)
Google Drive本地文件上传脚本
try: # 调用Drive v3 API file_metadata = {'name': 'file.jpg'} media = MediaFileUpload('file.jpg', mimetype='image/jpeg') # pylint: disable=maybe-no-member file = service.files().create(body=file_metadata, media_body=media, fields='id').execute() print(f'File ID: {file.get("id")}') except HttpError as error: print(f'An error occurred: {error}') file = None except Exception as e: print(e)
实现方案
Google Drive API提供的MediaIoBaseUpload类支持直接上传内存中的数据流,无需先写入本地文件。我们可以基于现有抓取逻辑,结合该类实现流式上传:
步骤1:导入必要模块
除原有的依赖外,补充导入内存流处理和Drive上传相关模块:
import io from googleapiclient.http import MediaIoBaseUpload
步骤2:改造数据流处理逻辑
将原本地保存的代码替换为Drive上传逻辑,针对不同内容类型处理数据流:
- HTML文本:将
response.text转为字节流,包装成io.BytesIO对象 - PDF文件:直接用
response.content包装成io.BytesIO - DOCX文件:利用
response.raw获取原始字节流,需重置流指针到起始位置
完整整合代码
以下是改造后的parse_page方法,整合了Drive上传逻辑(假设self.service是已初始化的Google Drive API服务对象):
def parse_page(self, link, q): response = requests.get(link, headers=self.headers, stream=True) print(f"\t- Response [{response.status_code}] for: {link}") published_date = self.get_date(response.text) # 确定文件类型、MIME类型和数据流 file_format = "" mime_type = "" stream = None if "text/html" in response.headers["content-type"]: file_format = ".html" mime_type = "text/html" # 将文本转为字节流 stream = io.BytesIO(response.text.encode('utf-8')) elif "application/pdf" in response.headers["content-type"]: file_format = ".pdf" mime_type = "application/pdf" stream = io.BytesIO(response.content) elif "officedocument.word" in response.headers["content-type"]: file_format = ".docx" mime_type = "application/vnd.openxmlformats-officedocument.wordprocessingml.document" # 使用原始响应流,重置指针到开头 stream = response.raw stream.seek(0) # 生成文件名 folder = urlparse(link).netloc filename = urlparse(link).path.split("/")[-1].replace("\n","") if filename == '': filename = folder + datetime.now().strftime("_%H_%M_%S") full_filename = filename + file_format # 上传到Google Drive try: file_metadata = {'name': full_filename} # 创建Media对象,启用分块上传(支持大文件断点续传) media = MediaIoBaseUpload(stream, mimetype=mime_type, chunksize=16*1024, resumable=True) # 调用API完成上传 file = self.service.files().create(body=file_metadata, media_body=media, fields='id').execute() print(f"\t- Uploaded to Drive, File ID: {file.get('id')}") except HttpError as error: print(f"\t- Drive upload error: {error}") except Exception as e: print(f"\t- Upload failed: {e}")
关键说明
- 分块上传:设置
chunksize=16*1024和resumable=True,既减少内存占用,又支持大文件断点续传 - 流指针重置:使用
response.raw时必须调用stream.seek(0),确保从流的起始位置读取数据 - MIME类型匹配:每个文件类型的MIME类型需与Google Drive兼容,避免文件识别错误
内容的提问来源于stack exchange,提问作者Mauro
相关产品推荐
相关产品推荐

