如何将Azure Blob的CSV文件以ANSI编码传输至SFTP服务器?
解决Azure Blob到SFTP的ANSI编码CSV传输问题
方法一:本地临时文件中转(最直观简单)
先将DataFrame生成本地ANSI编码的CSV文件,再通过Paramiko上传到SFTP,完全避开远程流操作的性能问题:
import pandas as pd from azure.storage.blob import BlobServiceClient import paramiko import os # 从Azure Blob读取数据到DataFrame azure_conn_str = "你的Azure存储连接字符串" container_name = "目标容器名" source_blob_name = "源CSV文件名.csv" blob_service_client = BlobServiceClient.from_connection_string(azure_conn_str) blob_client = blob_service_client.get_blob_client(container=container_name, blob=source_blob_name) df = pd.read_csv(blob_client.download_blob()) # 本地生成ANSI(cp1252)编码的临时CSV文件 temp_local_path = "temp_ansi_output.csv" df.to_csv(temp_local_path, encoding="cp1252", index=False) # SFTP上传临时文件 sftp_host = "SFTP服务器地址" sftp_user = "SFTP用户名" sftp_pass = "SFTP密码" remote_target_path = "/SFTP目标路径/target.csv" ssh_client = paramiko.SSHClient() ssh_client.set_missing_host_key_policy(paramiko.AutoAddPolicy()) ssh_client.connect(hostname=sftp_host, username=sftp_user, password=sftp_pass) with ssh_client.open_sftp() as sftp: sftp.put(temp_local_path, remote_target_path) # 清理资源 ssh_client.close() os.remove(temp_local_path)
方法二:内存字节流中转(无本地文件)
如果不想生成本地文件,可以在内存中直接生成ANSI编码的字节流,再上传到SFTP:
import pandas as pd from azure.storage.blob import BlobServiceClient import paramiko from io import StringIO, BytesIO # 读取Azure Blob数据 azure_conn_str = "你的Azure存储连接字符串" container_name = "目标容器名" source_blob_name = "源CSV文件名.csv" blob_service_client = BlobServiceClient.from_connection_string(azure_conn_str) blob_client = blob_service_client.get_blob_client(container=container_name, blob=source_blob_name) df = pd.read_csv(blob_client.download_blob()) # 在内存中生成ANSI编码的字节流 csv_str_io = StringIO() df.to_csv(csv_str_io, index=False) ansi_encoded_bytes = csv_str_io.getvalue().encode("cp1252") byte_stream = BytesIO(ansi_encoded_bytes) # SFTP上传字节流 sftp_host = "SFTP服务器地址" sftp_user = "SFTP用户名" sftp_pass = "SFTP密码" remote_target_path = "/SFTP目标路径/target.csv" ssh_client = paramiko.SSHClient() ssh_client.set_missing_host_key_policy(paramiko.AutoAddPolicy()) ssh_client.connect(hostname=sftp_host, username=sftp_user, password=sftp_pass) with ssh_client.open_sftp() as sftp: with sftp.file(remote_target_path, "wb") as remote_file: remote_file.write(byte_stream.getvalue()) ssh_client.close()
问题原因说明
直接通过SFTP远程文件对象调用df.to_csv()时,pandas的IO操作会频繁和远程SFTP流交互,而远程流的缓冲机制、网络延迟会导致写入效率极低甚至冻结。上述两种方法都是先在本地/内存完成编码和CSV序列化,再一次性上传,彻底避免了远程流的持续读写问题。
内容的提问来源于stack exchange,提问作者Mathieu
相关产品推荐
相关产品推荐

