读取ADLS中多JSON文件创建DataFrame时的循环覆盖问题
问题解决:避免DataFrame被循环覆盖
你的问题核心是双层嵌套循环导致重复赋值覆盖:外层循环遍历dataframes的每个名称,内层又遍历所有directories,所以每个DataFrame名称(json1、json2)都会先后被两个目录的文件数据赋值,最终只剩下最后一个目录的结果。
修正后的代码
# Primary storage info account_name = "account" # fill in your primary account name container_name = "json-logs" # fill in your container name subscription_id = "subs" resource_group = "resource-ML" # fill in your resource gropup for ADLS account workspace_name = "workspace-ML" # fill in your la workspace name dataframes = ['json1', 'json2'] directories=["test1","test2"] d = {} # 关键改动:用zip配对遍历,让每个DataFrame对应一个目录 for data, direct in zip(dataframes, directories): input_path = direct adls_path = f"abfss://{container_name}@{account_name}.dfs.core.windows.net/{input_path}" account_key = 'key' # Replace your storage account key new_path = input_path initialize_storage_account(account_name, account_key) pathlist = list_directory_contents(container_name, new_path, "json") input_file = pathlist[0].split("/")[-1] download_file_from_directory(container_name, new_path, input_file) json_normalize("output.json", "out_normalized.json") # 修正语法错误:直接用字典存储,无需globals() d[data] = pd.read_json('out_normalized.json') # 验证结果 print(d['json1'].head()) print(d['json2'].head())
核心改动说明
- 替换嵌套循环为配对遍历:用
zip(dataframes, directories)让json1对应test1目录,json2对应test2目录,每个DataFrame只会被赋值一次,彻底避免覆盖问题。 - 修正字典赋值语法:原来的
globals()d[data]是语法错误,直接用你定义的字典d存储DataFrame更规范,也方便后续统一管理多个DataFrame。 - 可选优化:每次处理完一个文件后,可添加代码删除临时的
output.json和out_normalized.json,避免下一次处理时读取到旧文件残留数据。
内容的提问来源于stack exchange,提问作者Lopa
相关产品推荐
相关产品推荐

