使用Python上传60MB+大文件到Blob存储的实现方法及问题排查
认知疑问解答
你的理解是正确的。你当前的实现逻辑是「浏览器向云函数提交文件的第三方下载链接、存储容器、目标文件名等少量参数」,文件从第三方站点下载、再上传到Blob存储的全流程都在云函数的运行环境中执行,全程走云服务的公网带宽,仅参数传输消耗本地极少量流量,完全适配你本地带宽不足的场景。
6-7KB上传异常问题排查
你代码里存在2个直接导致该问题的核心bug,还有3个可优化的风险点:
核心bug
- 参数获取逻辑遗漏了
link字段
你当前的参数提取逻辑只获取了container和file两个字段:
params = {key:req.params.get(key) for key in ("container", "file")}
但后续拼接下载地址时读取的是params.get("link"),这个值永远是None,最终请求的地址是https://www21.zippyshare.com/d/None,该地址返回的是站点的404/错误提示页面,大小刚好在6-7KB左右,你把这个错误页面当成了文件内容上传到Blob,才会出现只上传了几KB就结束、且没有报错的情况。
2. 下载请求未做合法性校验
你直接读取了请求返回的content内容,没有校验请求状态码、返回内容类型,即使请求返回的是错误页面,也会被当成正常文件内容上传,不会触发异常。
可优化的风险点
- 第三方站点反爬拦截:zippyshare这类公共文件分享站普遍有反爬策略,直接用requests默认请求头请求会被拦截,返回错误页面,需要补充模拟浏览器的
User-Agent、Referer等请求头。 - 大文件内存占用过高:你当前用
file.content会把整个文件加载到云函数的内存中,超过函数内存限制会被强制终止,60MB文件虽然不大,但更大的文件会直接触发OOM。可以结合stream=True特性,直接把请求返回的字节流传给Blob上传接口,无需全量加载到内存。 - 未做分块上传:超过一定大小的文件用单块上传容易失败,建议开启Blob的分块上传配置,提升大文件上传的成功率。
修复后参考代码
import requests import logging import azure.functions as func from azure.storage.blob import BlobServiceClient # 替换为你的Blob存储连接字符串 CONN = "你的Blob存储连接字符串" def download_file(link): # 补充模拟浏览器的请求头,绕过反爬 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36", "Referer": "https://www21.zippyshare.com/" } file = requests.get(link, stream=True, headers=headers, timeout=120) # 校验请求是否成功 file.raise_for_status() # 校验返回内容不是网页 if "text/html" in file.headers.get("Content-Type", ""): raise Exception("下载请求被拦截,返回的是网页内容") # 直接返回流对象,不加载到内存 return file.raw def get_blobclient(container=None, file_name=None): blob_service_client = BlobServiceClient.from_connection_string(CONN) blob_client = blob_service_client.get_blob_client(container, file_name) return blob_client def main(req: func.HttpRequest) -> func.HttpResponse: logging.info('Python HTTP trigger function processed a request.') error_message = None # 补充link参数的获取 params = {key:req.params.get(key) for key in ("container", "file", "link")} try: if not None in params.values(): link = f'https://www21.zippyshare.com/d/{params.get("link")}' blobClient = get_blobclient(container=params.get("container"), file_name=params.get("file")) logging.info("Uploading file...") # 流式上传,开启分块上传,允许覆盖 blobClient.upload_blob(download_file(link), blob_type="BlockBlob", overwrite=True, max_concurrency=4) logging.info("File has been uploaded") return func.HttpResponse( "File has been uploaded", status_code=200 ) else: error_message = f'{" ".join(str(k) for k,v in params.items() if v == None)}: These values have not been provided' logging.info(error_message) except Exception as error: logging.info(error) error_message = str(error) return func.HttpResponse( "Something went wrong..." if error_message is None else error_message, status_code=400 )
内容的提问来源于stack exchange,提问作者rarova
相关产品推荐
相关产品推荐

