You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

调用features_from_place()获取柏林数据时内存不足问题求助

解决features_from_place()返回GeoDataFrame列数过多导致的内存不足问题

首先注意你定义的tags字典存在重复键覆盖的问题——Python字典中同一个key多次赋值会保留最后一次的值,比如你的aeroway最终只会是'terminal',natural最终是'hill',public_transport最终是'station',这和你预期的多标签查询不符,先修正这个问题:

tags = {
    'aeroway': ['aerodrome', 'hangar', 'helipad', 'heliport', 'terminal'],
    'amenity': True,
    'building': True,
    'craft': True,
    'cuisine': True,
    'healthcare': True,
    'historic': True,
    'landuse': True,
    'leisure': True,
    'natural': ['beach', 'cave_entrance', 'hill'],
    'office': True,
    'public_transport': ['platform', 'station'],
    'shop': True,
    'sport': True,
    'tourism': True,
    'addr:street': True,
    'addr:flats': True
}

核心解决方案:限制返回的属性列

这类地理数据获取函数通常支持指定仅保留需要的属性字段,避免返回所有可能的OSM标签,这是减少列数、降低内存占用最直接的方法。

1. 明确指定需要保留的字段

以osmnx的features_from_place为例,可通过attributes参数指定要保留的列,只保留你实际需要的标签:

import osmnx as ox

# 定义需要保留的核心字段(geometry为地理数据必需,加上你业务需要的标签)
keep_attrs = ['geometry', 'name', 'aeroway', 'amenity', 'building', 'natural', 'public_transport', 'addr:street', 'addr:flats']

# 获取数据时指定仅保留这些字段
gdf = ox.features_from_place(
    {"city": "Berlin","country": "Germany"},
    tags=tags,
    attributes=keep_attrs
)

这样返回的GeoDataFrame只会包含你指定的列,列数会大幅缩减,内存占用直接降低。

2. 拆分查询后统一列结构再合并

如果必须拆分tags多次查询,可通过统一列结构解决列数不同无法合并的问题:

  • 先定义所有需要保留的列集合
  • 每次查询后,只保留这些列,缺失的列填充NaN
  • 最后合并所有GeoDataFrame

示例代码:

import pandas as pd
import osmnx as ox

# 拆分后的tags分组
tag_groups = [
    {'aeroway': ['aerodrome', 'hangar', 'helipad', 'heliport', 'terminal']},
    {'amenity': True, 'building': True},
    {'natural': ['beach', 'cave_entrance', 'hill']},
    # 其他需要拆分的tag组...
]

# 定义统一要保留的列
target_cols = ['geometry', 'name', 'aeroway', 'amenity', 'building', 'natural', 'public_transport', 'addr:street', 'addr:flats']

gdf_list = []
for tag_group in tag_groups:
    temp_gdf = ox.features_from_place({"city": "Berlin","country": "Germany"}, tags=tag_group)
    # 只保留目标列,缺失列补NaN
    temp_gdf = temp_gdf.reindex(columns=target_cols)
    gdf_list.append(temp_gdf)

# 合并所有数据
final_gdf = pd.concat(gdf_list, ignore_index=True)

3. 优化数据类型进一步压缩内存

即使列数减少,还可以通过转换数据类型降低内存占用:

  • 将重复值较多的object类型列转为category类型
  • 对数值型列(如果有)转为更小的数值类型(如int32代替int64)

示例:

# 将非几何的object列转为category类型
for col in final_gdf.columns:
    if col != 'geometry' and final_gdf[col].dtype == 'object':
        final_gdf[col] = final_gdf[col].astype('category')

4. 过滤不必要的要素

如果某些标签对应的要素对你的分析没有价值,可以直接从tags中移除,减少返回的行数和列数——比如如果craft、cuisine这类标签没用,就删掉它们,进一步缩小数据规模。

内容的提问来源于stack exchange,提问作者Blad To

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 08:36:00