ADF传参至Databricks Notebook的每日数据加载异常问题
解决方案
方案一:Pipeline层面拆分清空与写入操作(推荐)
把清空当日文件夹的逻辑从Notebook中剥离,放到ForEach循环之前作为独立Activity执行,这样只会清空一次,后续循环仅负责写入数据:
在Azure Data Factory/Synapse Pipeline中,添加一个Delete Activity,配置要删除的路径为:
/mnt/bronze/attribute_code/@pipeline().parameters.day开启递归删除(Recursive)选项。
修改Notebook代码,移除删除/创建文件夹的逻辑,同时调整文件名避免重复覆盖:
data, response = get_data_url(url=f"https://p.cloud.com/api/rest/v1/attributes/{attribute_code}/options",access_token=access_token) # 确保文件夹存在(可选,Delete后写入会自动创建父目录) dbutils.fs.mkdirs(f'/mnt/bronze/attribute_code/{day}') # 用attribute_code作为文件名的一部分,避免重复覆盖 dbutils.fs.put(f'/mnt/bronze/attribute_code/{day}/data_{attribute_code}.json', response.text)若需要按顺序命名,也可结合时间戳生成唯一文件名:
from datetime import datetime timestamp = datetime.now().strftime("%Y%m%d%H%M%S%f") dbutils.fs.put(f'/mnt/bronze/attribute_code/{day}/data_{timestamp}.json', response.text)
方案二:Notebook内部实现单次清空逻辑
如果无法修改Pipeline结构,可以在Notebook中通过标记文件实现仅首次执行时清空文件夹:
from datetime import datetime data, response = get_data_url(url=f"https://p.cloud.com/api/rest/v1/attributes/{attribute_code}/options",access_token=access_token) folder_path = f'/mnt/bronze/attribute_code/{day}' flag_file = f'{folder_path}/_initialized' # 仅当标记文件不存在时,执行清空操作 if not dbutils.fs.exists(flag_file): dbutils.fs.rm(folder_path, True) dbutils.fs.mkdirs(folder_path) # 创建标记文件,标记已完成初始化 dbutils.fs.put(flag_file, f"Initialized at {datetime.now()}", overwrite=True) # 生成唯一文件名,避免覆盖 dbutils.fs.put(f'{folder_path}/data_{attribute_code}.json', response.text)
关键说明
- 方案一优势:逻辑清晰,避免多实例竞争,性能更优,适合绝大多数场景。
- 方案二注意:如果ForEach循环是并行执行,可能存在多个Notebook实例同时检查标记文件的情况,常规场景下标记文件方式足以应对,若需要更严谨可添加短暂延迟或使用Azure Blob Lease实现分布式锁。
- 文件名唯一化:必须替换原固定的
count=0,否则会导致不同attribute_code的数据互相覆盖,推荐使用attribute_code或时间戳作为文件名的一部分。
内容的提问来源于stack exchange,提问作者Greencolor
相关产品推荐
相关产品推荐

