Azure Automation中Python脚本参数化对接Data Lake Gen2生成JSON Schema
改造JSON Schema生成代码适配Azure Data Lake Gen2与Azure Automation
改造目标
将现有本地文件的JSON Schema生成代码,封装为可接收Azure Data Lake Gen2(ADLS Gen2)路径参数的模块,适配Azure Automation运行环境。
依赖准备
首先需要安装必要的Python包:
pip install genson azure-storage-file-datalake
如果使用Azure Automation托管身份认证,还需安装:
pip install azure-identity
改造后的完整代码
import json from genson import SchemaBuilder from azure.storage.filedatalake import DataLakeServiceClient from azure.core.exceptions import ResourceNotFoundError from azure.identity import DefaultAzureCredential # 托管身份认证时使用 def generate_json_schema(adls_account_name, input_file_path, output_file_path, adls_account_key=None): # 初始化ADLS服务客户端(支持密钥或托管身份认证) if adls_account_key: service_client = DataLakeServiceClient( account_url=f"https://{adls_account_name}.dfs.core.windows.net", credential=adls_account_key ) else: # 使用Azure Automation托管身份认证 service_client = DataLakeServiceClient( account_url=f"https://{adls_account_name}.dfs.core.windows.net", credential=DefaultAzureCredential() ) # 读取ADLS中的源JSON文件 try: fs_name, file_path = input_file_path.split('/', 1) file_client = service_client.get_file_client(fs_name, file_path) download_stream = file_client.download_file() json_content = json.loads(download_stream.readall().decode('utf-8')) except ResourceNotFoundError: raise ValueError(f"输入文件不存在:{input_file_path}") except Exception as e: raise RuntimeError(f"读取ADLS文件失败:{str(e)}") # 生成JSON Schema schema_builder = SchemaBuilder() schema_builder.add_object(json_content) schema_output = schema_builder.to_json(indent=2) # 将Schema写入ADLS目标路径 try: output_fs_name, output_file_full_path = output_file_path.split('/', 1) output_file_client = service_client.get_file_client(output_fs_name, output_file_full_path) # 覆盖已存在的文件(可选,根据需求调整) try: output_file_client.delete_file() except ResourceNotFoundError: pass output_file_client.upload_data(schema_output, overwrite=True) except Exception as e: raise RuntimeError(f"写入ADLS文件失败:{str(e)}") if __name__ == "__main__": # 示例参数(在Azure Automation中可通过Runbook输入参数传递) ADLS_ACCOUNT = "your_adls_account_name" INPUT_PATH = "raw-data/supplier-data/test.json" # 格式:文件系统名/文件路径 OUTPUT_PATH = "schema-store/supplier-schemas/test.schema.json" # ADLS_ACCOUNT_KEY = "your_adls_account_key" # 使用密钥认证时取消注释 # 调用函数(使用托管身份时无需传密钥) generate_json_schema(ADLS_ACCOUNT, INPUT_PATH, OUTPUT_PATH) # generate_json_schema(ADLS_ACCOUNT, INPUT_PATH, OUTPUT_PATH, ADLS_ACCOUNT_KEY) # 密钥认证方式
关键说明
- ADLS路径格式:输入输出路径需遵循
文件系统名/文件相对路径格式,例如raw-data/invoices/2024.json - 认证方式:
- 密钥认证:直接传入存储账号密钥,适合测试场景
- 托管身份认证:使用
DefaultAzureCredential,无需硬编码密钥,符合Azure生产环境安全规范,需确保Automation账户的托管身份拥有ADLS的Storage Blob Data Contributor或对应权限
- Azure Automation适配:
- 将ADLS账号名、输入输出路径设置为Runbook的输入参数,替换代码中的示例值
- 在Automation账户中导入所需的Python包(genson、azure-storage-file-datalake、azure-identity)
- 错误处理:包含文件不存在、读写失败等异常捕获,便于在Automation中排查问题
内容的提问来源于stack exchange,提问作者paone
相关产品推荐
相关产品推荐

