如何实现requests流式下载CSV至GCP Cloud Storage的流式上传?
问题:流式将API返回的CSV上传至GCP Cloud Storage
需求背景
- 通过
requests.get获取API输出的CSV文件 - 希望流式上传到GCP Cloud Storage,避免将完整文件下载到本地占用存储
尝试的代码
import requests from google.cloud import storage url = "https://people.sc.fsu.edu/~jburkardt/data/csv/addresses.csv" # GCP信息 client = storage.Client(project="my-project") bucket = client.get_bucket('my-bucket') target_blob = bucket.blob("test/report_01.csv") with requests.get(url, stream=True) as f: target_blob.upload_from_file(f)
报错信息
AttributeError: 'Response' object has no attribute 'tell'
补充说明
- 已知有类似问题,但现有解决方案是先读取完整文件再上传,不符合流式需求
解决思路
方法1:包装Response对象适配GCS要求
GCS的upload_from_file要求传入的文件对象支持tell()方法(用于追踪上传进度),但requests开启stream=True后的Response对象没有这个方法。我们可以给Response做个简单包装,补上这个方法:
import requests from google.cloud import storage class StreamableResponse: def __init__(self, response): self.response = response self.position = 0 def read(self, size=None): chunk = self.response.read(size) self.position += len(chunk) return chunk def tell(self): return self.position url = "https://people.sc.fsu.edu/~jburkardt/data/csv/addresses.csv" # GCP信息 client = storage.Client(project="my-project") bucket = client.get_bucket('my-bucket') target_blob = bucket.blob("test/report_01.csv") with requests.get(url, stream=True) as r: r.raise_for_status() # 确保请求无错误 streamable = StreamableResponse(r) # 可根据网络情况调整chunk_size,比如1MB或更大 target_blob.upload_from_file(streamable, chunk_size=1024*1024)
方法2:使用可恢复上传API实现流式传输
如果需要更可靠的流式上传(比如大文件),可以用GCS的可恢复上传接口,分块传输内容:
import requests from google.cloud import storage from google.resumable_media.requests import ResumableUpload url = "https://people.sc.fsu.edu/~jburkardt/data/csv/addresses.csv" # GCP信息 client = storage.Client(project="my-project") bucket = client.get_bucket('my-bucket') target_blob = bucket.blob("test/report_01.csv") # 初始化可恢复上传会话 upload_url = target_blob.create_resumable_upload_session() upload = ResumableUpload(upload_url, chunk_size=1024*1024) with requests.get(url, stream=True) as r: r.raise_for_status() # 迭代读取响应分块,逐步上传 for chunk in r.iter_content(chunk_size=1024*1024): if chunk: upload.transmit_next_chunk(chunk) # 完成上传 upload.finish()
注意事项
- 两种方法都实现了边下载边上传,不会在本地存储完整文件
chunk_size建议设置为1MB~10MB,太小会增加请求次数,太大会占用更多内存- 务必添加
r.raise_for_status()处理请求错误,避免上传损坏的内容
内容的提问来源于stack exchange,提问作者Ralubrusto
相关产品推荐
相关产品推荐

