You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何优化处理大型GeoJSON地理数据文件,提取指定字段生成目标结构

大型GeoJSON指定字段提取方案

针对你描述的标准GeoJSON格式文件,最优方案根据文件大小分为两种:

小文件方案(≤1GB,内存可容纳全量数据)

直接用Python内置json库全量读取处理,代码实现最简单:

import json

# 替换为你的源文件路径
with open("geo_data.json", "r", encoding="utf-8") as f:
    raw_data = json.load(f)

# 如果你需要所有数据合并到一个根对象,用这个结构
output = {
    "Number": [],
    "Area sq.m.": [],
    "Forest cov": "None",
    "Status": "None"
}

for feature in raw_data["features"]:
    props = feature["properties"]
    output["Number"].append(props["id"])
    output["Area sq.m."].append(props["area_value"])

# 输出结果到文件
with open("output.json", "w", encoding="utf-8") as f:
    json.dump(output, f, ensure_ascii=False, indent=2)

大文件方案(≥1GB,全量加载会内存溢出)

用第三方流式解析库ijson逐行读取features数组的元素,不会占用过多内存:

  1. 先安装依赖:pip install ijson
  2. 处理代码:
import ijson
import json

output = {
    "Number": [],
    "Area sq.m.": [],
    "Forest cov": "None",
    "Status": "None"
}

with open("large_geo_data.json", "r", encoding="utf-8") as f:
    # 流式遍历features下的每个元素
    for feature in ijson.items(f, "features.item"):
        props = feature["properties"]
        output["Number"].append(props["id"])
        output["Area sq.m."].append(props["area_value"])

with open("output.json", "w", encoding="utf-8") as f:
    json.dump(output, f, ensure_ascii=False, indent=2)

补充说明

如果你的需求是每个id对应一条独立的结构对象,只需要把上面的output结构调整为列表,每次遍历append单个对象即可:

output = []
for feature in raw_data["features"]:
    props = feature["properties"]
    item = {
        "Number": [props["id"]],
        "Area sq.m.": [props["area_value"]],
        "Forest cov": "None",
        "Status": "None"
    }
    output.append(item)

如果存在部分数据缺失id或area_value字段,可在遍历逻辑里添加判空跳过/默认值填充逻辑,避免运行报错。

内容的提问来源于stack exchange,提问作者Go_ Dmt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 09:57:02