Databricks中追加写入Azure Blob文件报错OSError: [Errno 95]的问题及解决方案咨询
Databricks中追加写入Azure Blob文件报错OSError: [Errno 95]的问题及解决方案咨询
嗨,我来帮你拆解这个问题~
首先得搞清楚为什么用open('\mnt\file\','a')会报错,而'w'模式却能成功:
这是因为Databricks挂载的Azure Blob默认是块Blob类型,它的底层存储机制决定了不支持直接追加写入操作。块Blob是由多个独立的块组成的,一旦块上传完成就只能被替换或删除,没法直接在现有文件末尾追加内容。而'w'模式是直接创建新文件(如果原文件存在就覆盖),相当于重新上传一个全新的块Blob,所以能正常执行。
接下来给你几个实用的 workaround,按需选择:
用Databricks原生工具dbutils.fs模拟追加
如果文件不大,可以先读取现有文件的内容,把新内容拼接进去后再重新写入。代码示例:# 读取现有文件内容(如果文件很大不建议用这个方法) existing_content = dbutils.fs.head("/mnt/your_file_path") # 准备要追加的内容 new_content = "需要追加的内容\n" # 合并后写入(overwrite=True会覆盖原文件,等同于模拟追加) dbutils.fs.put("/mnt/your_file_path", existing_content + new_content, overwrite=True)改用Azure Blob的追加Blob类型
如果你的场景是频繁追加内容(比如日志记录),可以把目标Blob创建为追加Blob(这是Azure Blob专门为追加场景设计的类型),然后用Azure Storage SDK直接操作。代码示例:from azure.storage.blob import BlobServiceClient # 替换成你的存储连接字符串 conn_str = "your_azure_storage_connection_string" blob_service_client = BlobServiceClient.from_connection_string(conn_str) # 获取追加Blob的客户端 append_blob_client = blob_service_client.get_append_blob_client(container="your_container", blob="your_target_blob") # 如果Blob不存在,先创建它 if not append_blob_client.exists(): append_blob_client.create_append_blob() # 执行追加操作 append_blob_client.append_block("要追加的内容\n")用Delta Lake处理结构化数据的追加
如果是处理结构化数据(比如CSV、Parquet),更推荐用Databricks原生的Delta Lake,它支持ACID事务,原生支持追加写入,还能避免数据一致性问题。代码示例:# 读取现有Delta表(如果是第一次就跳过这步,直接写) existing_df = spark.read.format("delta").load("/mnt/delta_table_path") # 准备要追加的新数据DataFrame new_data_df = spark.createDataFrame([(1, "new_record"), (2, "another_record")], ["id", "content"]) # 追加写入Delta表 new_data_df.write.format("delta").mode("append").save("/mnt/delta_table_path")
备注:内容来源于stack exchange,提问作者sam yi
相关产品推荐
相关产品推荐

