Python实现Google Cloud Storage可恢复上传异常及弱网小文件方案咨询
GCS可恢复上传测试问题及小文件断点续传方案
500MB文件可恢复上传测试异常原因
你当前的测试方式无法验证可恢复上传效果:删除文件后重新上传相当于发起全新的上传会话,之前中断的进度无法复用,因此会从头开始上传,耗时与首次一致,这不是可恢复上传的正常表现。
正确的测试流程:
- 保留GCS上的目标文件(可恢复上传会话与目标对象绑定)
- 上传中断后,复用同一个
Blob对象调用上传方法,客户端库会自动检测并恢复之前的上传会话,继续未完成的部分。
另外,若代码每次上传都重新创建storage.Client和Blob对象,也可能导致会话信息丢失,无法恢复进度。
小文件(0.2-1MB)的可恢复上传方案
由于GCS官方对小于8MB的文件默认使用不可恢复的分块上传,针对树莓派+不稳定蜂窝网络的场景,推荐以下两种可行方案:
方案1:手动分块上传+合并
将小文件拆分为更小的固定大小块(比如64KB),逐个上传为临时对象,记录已完成的块,网络恢复后仅上传未完成的部分,最后合并为完整文件:
import os from google.cloud import storage def split_file(file_path, chunk_size=64*1024): chunks = [] with open(file_path, 'rb') as f: while True: chunk = f.read(chunk_size) if not chunk: break chunks.append(chunk) return chunks def upload_small_file_resumable(blob_name, file_path, bucket_name, credentials_path): storage_client = storage.Client.from_service_account_json(credentials_path) bucket = storage_client.bucket(bucket_name) chunks = split_file(file_path) chunk_blob_names = [f"{blob_name}.chunk.{i}" for i in range(len(chunks))] completed_chunks = [] # 检查已上传的块 for idx, chunk_blob_name in enumerate(chunk_blob_names): if bucket.blob(chunk_blob_name).exists(): completed_chunks.append(idx) # 上传未完成的块 for idx, chunk in enumerate(chunks): if idx not in completed_chunks: chunk_blob = bucket.blob(chunk_blob_names[idx]) chunk_blob.upload_from_string(chunk) # 合并所有块为最终文件 if len(completed_chunks) == len(chunks): source_blobs = [bucket.blob(name) for name in chunk_blob_names] bucket.blob(blob_name).compose(source_blobs) # 删除临时块 for blob in source_blobs: blob.delete()
方案2:强制启用可恢复上传会话
手动创建可恢复上传会话,强制对小文件使用可恢复上传逻辑,记录会话URI和已上传字节数,中断后从断点继续:
from google.cloud import storage import os def upload_small_file_resumable_session(blob_name, file_path, bucket_name, credentials_path): storage_client = storage.Client.from_service_account_json(credentials_path) bucket = storage_client.bucket(bucket_name) blob = bucket.blob(blob_name) file_size = os.path.getsize(file_path) session_uri_path = f"{blob_name}.session.txt" uploaded_bytes = 0 # 检查是否有保存的会话和进度 if os.path.exists(session_uri_path): with open(session_uri_path, 'r') as f: session_uri, uploaded_bytes = f.read().split(',') uploaded_bytes = int(uploaded_bytes) else: # 创建新的可恢复上传会话 session_uri = blob.create_resumable_upload_session(content_type='application/octet-stream') with open(session_uri_path, 'w') as f: f.write(f"{session_uri},{uploaded_bytes}") # 从断点开始上传 with open(file_path, 'rb') as f: f.seek(uploaded_bytes) remaining_data = f.read() if remaining_data: response = blob.resumable_upload(session_uri, remaining_data, start_byte=uploaded_bytes) uploaded_bytes += len(remaining_data) # 更新进度 with open(session_uri_path, 'w') as f: f.write(f"{session_uri},{uploaded_bytes}") # 上传完成后删除会话记录 if uploaded_bytes == file_size: if os.path.exists(session_uri_path): os.remove(session_uri_path)
注意事项:
- 方案2需在本地保存会话URI和上传进度(如文本文件),树莓派断电后也可恢复
- 两种方案都需添加网络异常捕获与重试逻辑
- 方案1需确保临时块命名唯一,避免冲突
内容的提问来源于stack exchange,提问作者Richard
相关产品推荐
相关产品推荐

