如何通过ADF或Notebook将JSON文件数据转换为Base64二进制格式?
解决方案:将JSON文件转换为Base64格式
一、Azure Data Factory(ADF)适用场景(小文件)
如果你的JSON文件体积较小(通常<4MB),可以通过以下步骤实现:
- 配置Lookup活动,源选择你的JSON数据集,在活动设置中勾选「读取整个内容」(Read entire content)。
- 使用表达式
@base64(body('Lookup'))将读取到的文件内容转换为Base64格式。 - 可将转换后的结果存入变量、写入存储或后续流程使用。
注意:Lookup活动对读取的文件大小有限制,大文件建议使用Databricks处理。
二、Databricks Notebook实现代码(支持大文件)
由于你已将Blob存储挂载到Databricks,直接使用以下代码即可完成完整文件的Base64转换:
import base64 # 替换为你的挂载路径下的JSON文件地址 source_json_path = "/dbfs/mnt/your_mount_name/path/to/target.json" # 替换为Base64结果的输出路径 output_base64_path = "/dbfs/mnt/your_mount_name/path/to/output_base64.txt" # 读取文件二进制内容 with open(source_json_path, "rb") as json_file: binary_data = json_file.read() # 转换为Base64字符串 base64_result = base64.b64encode(binary_data).decode("utf-8") # 将结果写入文件 with open(output_base64_path, "w") as output_file: output_file.write(base64_result) # 可选:验证结果(仅打印前100个字符) print("Base64转换结果预览:", base64_result[:100] + "...")
代码说明
- 使用
rb模式读取文件,确保获取完整的二进制数据,避免文本编码导致的内容丢失。 base64.b64encode返回字节流,需用decode("utf-8")转换为可存储的字符串格式。- 若处理超大文件(GB级),可采用分块读取转换的方式,避免内存溢出:
import base64 chunk_size = 1024 * 1024 # 1MB分块 source_json_path = "/dbfs/mnt/your_mount_name/path/to/large_file.json" output_base64_path = "/dbfs/mnt/your_mount_name/path/to/large_output_base64.txt" with open(source_json_path, "rb") as json_file, open(output_base64_path, "wb") as output_file: while chunk := json_file.read(chunk_size): output_file.write(base64.b64encode(chunk))
常见问题排查
如果你的原有代码无法工作,可能是以下原因:
- 路径错误:使用Python原生
open时需添加/dbfs前缀(挂载路径的完整路径),而使用dbutils.fs时无需该前缀。 - 读取模式错误:误用
r文本模式读取,导致二进制内容被转码损坏。 - 编码问题:转换Base64时未正确将字节流解码为字符串。
内容的提问来源于stack exchange,提问作者Nezko1
相关产品推荐
相关产品推荐

