在Azure机器学习环境中读取Azure Blob存储中的Shapefile
在Azure机器学习中读取Blob存储中的Shapefile
前置条件
- 环境已安装
geopandas与fiona(若未安装,可在脚本开头执行!pip install geopandas fiona完成安装) - 已获取Azure ML工作区的
datastore对象(你此前读取Parquet/CSV时已用到该对象)
实现步骤
1. 下载Shapefile配套文件至本地临时目录
Shapefile依赖.shp、.dbf、.shx、.prj等同名配套文件才能正常读取,需将同一Shapefile的所有关联文件下载至同一本地目录:
from azureml.core import Workspace, Datastore import os # 加载工作区(若已存在ws对象可跳过此步) ws = Workspace.from_config() # 指定目标datastore datastore = Datastore.get(ws, datastore_name="你的datastore名称") # Blob中Shapefile的前缀(例:所有配套文件均为"shapefiles/region_data/area_boundary.*") shapefile_prefix = "shapefiles/region_data/area_boundary" # 创建本地临时存储目录 local_temp_dir = "./temp_shapefile" os.makedirs(local_temp_dir, exist_ok=True) # 下载所有匹配前缀的文件 datastore.download(target_path=local_temp_dir, prefix=shapefile_prefix, overwrite=True)
2. 读取Shapefile数据
使用geopandas读取本地下载好的.shp文件:
import geopandas as gpd # 构造本地.shp文件路径 local_shp_path = os.path.join(local_temp_dir, f"{os.path.basename(shapefile_prefix)}.shp") # 读取Shapefile为GeoDataFrame geo_df = gpd.read_file(local_shp_path) # 验证读取结果 print(geo_df.head())
批量读取多个Shapefile的方案
若Blob中存在多个独立的Shapefile(每个Shapefile对应一个子目录),可通过遍历目录批量处理:
# 列出Blob中目标目录下的所有子目录(每个子目录对应一个Shapefile) shapefile_subdirs = datastore.ls(path="shapefiles/") for subdir in shapefile_subdirs: # 下载当前子目录下的所有文件 datastore.download(target_path="./temp_shapefiles", prefix=subdir, overwrite=True) # 定位子目录中的.shp文件 shp_file_list = [f for f in os.listdir(os.path.join("./temp_shapefiles", subdir)) if f.endswith(".shp")] if shp_file_list: target_shp_path = os.path.join("./temp_shapefiles", subdir, shp_file_list[0]) geo_df = gpd.read_file(target_shp_path) # 在此添加数据处理逻辑(如合并、分析等) print(f"完成读取:{subdir}")
注意事项
- 确保同一Shapefile的所有配套文件文件名完全一致,否则会读取失败
- AML计算实例的临时目录(如
/tmp)会在实例重启后清空,若需持久化数据,读完后可上传回Blob或保存至AML数据集 - 若使用AML管道,建议使用
tempfile模块创建临时目录,避免目录冲突
内容的提问来源于stack exchange,提问作者Vanaclocha
相关产品推荐
相关产品推荐

