调用features_from_place()获取柏林数据时内存不足问题求助
解决features_from_place()返回GeoDataFrame列数过多导致的内存不足问题
首先注意你定义的tags字典存在重复键覆盖的问题——Python字典中同一个key多次赋值会保留最后一次的值,比如你的aeroway最终只会是'terminal',natural最终是'hill',public_transport最终是'station',这和你预期的多标签查询不符,先修正这个问题:
tags = { 'aeroway': ['aerodrome', 'hangar', 'helipad', 'heliport', 'terminal'], 'amenity': True, 'building': True, 'craft': True, 'cuisine': True, 'healthcare': True, 'historic': True, 'landuse': True, 'leisure': True, 'natural': ['beach', 'cave_entrance', 'hill'], 'office': True, 'public_transport': ['platform', 'station'], 'shop': True, 'sport': True, 'tourism': True, 'addr:street': True, 'addr:flats': True }
核心解决方案:限制返回的属性列
这类地理数据获取函数通常支持指定仅保留需要的属性字段,避免返回所有可能的OSM标签,这是减少列数、降低内存占用最直接的方法。
1. 明确指定需要保留的字段
以osmnx的features_from_place为例,可通过attributes参数指定要保留的列,只保留你实际需要的标签:
import osmnx as ox # 定义需要保留的核心字段(geometry为地理数据必需,加上你业务需要的标签) keep_attrs = ['geometry', 'name', 'aeroway', 'amenity', 'building', 'natural', 'public_transport', 'addr:street', 'addr:flats'] # 获取数据时指定仅保留这些字段 gdf = ox.features_from_place( {"city": "Berlin","country": "Germany"}, tags=tags, attributes=keep_attrs )
这样返回的GeoDataFrame只会包含你指定的列,列数会大幅缩减,内存占用直接降低。
2. 拆分查询后统一列结构再合并
如果必须拆分tags多次查询,可通过统一列结构解决列数不同无法合并的问题:
- 先定义所有需要保留的列集合
- 每次查询后,只保留这些列,缺失的列填充
NaN - 最后合并所有GeoDataFrame
示例代码:
import pandas as pd import osmnx as ox # 拆分后的tags分组 tag_groups = [ {'aeroway': ['aerodrome', 'hangar', 'helipad', 'heliport', 'terminal']}, {'amenity': True, 'building': True}, {'natural': ['beach', 'cave_entrance', 'hill']}, # 其他需要拆分的tag组... ] # 定义统一要保留的列 target_cols = ['geometry', 'name', 'aeroway', 'amenity', 'building', 'natural', 'public_transport', 'addr:street', 'addr:flats'] gdf_list = [] for tag_group in tag_groups: temp_gdf = ox.features_from_place({"city": "Berlin","country": "Germany"}, tags=tag_group) # 只保留目标列,缺失列补NaN temp_gdf = temp_gdf.reindex(columns=target_cols) gdf_list.append(temp_gdf) # 合并所有数据 final_gdf = pd.concat(gdf_list, ignore_index=True)
3. 优化数据类型进一步压缩内存
即使列数减少,还可以通过转换数据类型降低内存占用:
- 将重复值较多的
object类型列转为category类型 - 对数值型列(如果有)转为更小的数值类型(如
int32代替int64)
示例:
# 将非几何的object列转为category类型 for col in final_gdf.columns: if col != 'geometry' and final_gdf[col].dtype == 'object': final_gdf[col] = final_gdf[col].astype('category')
4. 过滤不必要的要素
如果某些标签对应的要素对你的分析没有价值,可以直接从tags中移除,减少返回的行数和列数——比如如果craft、cuisine这类标签没用,就删掉它们,进一步缩小数据规模。
内容的提问来源于stack exchange,提问作者Blad To
相关产品推荐
相关产品推荐

