Azure机器学习:如何读取文件夹型数据资产中的单个文件
解决Azure ML中读取文件夹数据资产单个文件的问题
错误原因
你错误地使用folder参数传入单个文件路径,mltable会将其当作文件夹路径去查找,导致找不到目标文件。
正确解决方案
方法一:直接指定单个文件路径(使用file参数)
将路径参数从folder改为file,同时确保文件路径拼接正确:
import mltable from azure.ai.ml import MLClient from azure.identity import DefaultAzureCredential ml_client = MLClient.from_config(credential=DefaultAzureCredential()) data_asset = ml_client.data.get("data_asset_name", version="1") # 确保基础路径末尾带斜杠,避免拼接路径出错 base_path = data_asset.path if data_asset.path.endswith("/") else data_asset.path + "/" file_full_path = base_path + "file_name.csv" # 使用file参数指定单个文件 path_config = { 'file': file_full_path } tbl = mltable.from_delimited_files(paths=[path_config]) df = tbl.to_pandas_dataframe() df
方法二:通过glob筛选单个文件(基于文件夹资产)
如果需要保留文件夹资产的上下文,可以用glob_pattern精准匹配目标文件:
import mltable from azure.ai.ml import MLClient from azure.identity import DefaultAzureCredential ml_client = MLClient.from_config(credential=DefaultAzureCredential()) data_asset = ml_client.data.get("data_asset_name", version="1") path_config = { 'folder': data_asset.path } # 用glob_pattern匹配单个文件 tbl = mltable.from_delimited_files( paths=[path_config], glob_pattern="file_name.csv" ) df = tbl.to_pandas_dataframe() df
内容的提问来源于stack exchange,提问作者Ameya Bhave
相关产品推荐
相关产品推荐

