You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python将Azure Blob存储的CSV文件转换为Excel并回传同一容器

实现方案

首先需要安装xlsx导出依赖:

pip install openpyxl

pandas导出xlsx格式需要依赖openpyxl作为处理引擎,不安装会触发模块缺失报错。


完整可运行代码

from io import StringIO, BytesIO
from azure.storage.blob import BlobServiceClient, ContainerClient, BlobClient
from typing import Container
import pandas as pd

# 替换为自己的存储账户连接字符串
# 国际版Azure EndpointSuffix为core.windows.net,中国区Azure为core.chinacloudapi.cn
conn_str = "DefaultEndpointsProtocol=https;AccountName=你的账户名;AccountKey=你的账户密钥;EndpointSuffix=core.chinacloudapi.cn"
container = "testing"
blob_name = "Test.csv"
filename = "test.xlsx"

container_client = ContainerClient.from_connection_string(
    conn_str=conn_str, 
    container_name=container
)   

# 原有读取CSV逻辑
downloaded_blob = container_client.download_blob(blob_name)
read_file = pd.read_csv(StringIO(downloaded_blob.content_as_text()))

# 将DataFrame转为内存中的Excel字节流
output = BytesIO()
# index=False表示不导出pandas默认索引列,可根据业务需求调整
read_file.to_excel(output, engine='openpyxl', index=False)
# 重置流指针到起始位置,避免上传内容为空
output.seek(0)

# 上传Excel文件到同一存储容器
container_client.upload_blob(
    name=filename,
    data=output,
    overwrite=True # 不需要覆盖已有文件可删除该参数
)

注意事项

  • 连接字符串请替换为自身存储账户的真实信息,账户密钥需要有容器的读写权限
  • 若处理的CSV文件体积超过可用内存,建议采用分批读取、分批写入的逻辑,避免内存溢出
  • 有自定义Excel格式需求时,可以调用openpyxl的接口对生成的字节流做样式调整后再上传

内容的提问来源于stack exchange,提问作者Sreepu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 01:45:03