You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python解析GeoJSON转DataFrame时出现数据行数缺失问题

问题描述

我正尝试从市级行政区域GeoJSON文件中提取韩国各自治市的韩文、英文名称,所用Python代码如下:

import json 
import pandas as pd

Korean_municipalities = json.load(open('skorea-municipalities-2018-geo.json', 'r'))
munic_map_eng = {}

for feature in Korean_municipalities['features']:
        feature['id'] = feature['properties']['name_eng']
        munic_map_eng[feature['properties']['name']] = feature['id']  


df_munic = pd.DataFrame(list(munic_map_eng.items()))  

运行后出现数据缺失问题:

  • 统计得len(Korean_municipalities['features']) = 250,即文件内共包含250个市级行政单元,但最终生成的DataFrame对象df_munic维度为df_munic.shape = (227,2),共缺失23条记录。
  • 将相同逻辑代码用于省级、sub-municipality级行政区域名称提取时:
    • sub-municipality级共3504个行政单元,最终生成的DataFrame仅3142行,存在同类缺失问题
    • 省级共17个行政单元,提取结果完全正常
问题原因

核心原因是Python字典的键唯一性规则:
你用韩文行政区名feature['properties']['name']作为字典的键,韩国省级行政区共17个,本身没有重名,所以不会出现键覆盖,结果完全正常。但市级、sub-municipality级行政区存在大量跨上级区域的重名情况,遍历feature时,后出现的重名条目会直接覆盖字典中已存在的同键值,最终字典条目数必然小于实际feature总数,转成DataFrame就会出现记录缺失。

修正方案

不要用字典做中间存储触发去重逻辑,直接遍历所有feature时将需要的字段存入列表,再转成DataFrame即可避免覆盖问题,参考代码:

import json 
import pandas as pd

Korean_municipalities = json.load(open('skorea-municipalities-2018-geo.json', 'r'))
record_list = []

for feature in Korean_municipalities['features']:
    # 按需提取字段,可加入行政编码、上级行政区名等字段做唯一区分
    record_list.append({
        "name_korean": feature['properties']['name'],
        "name_english": feature['properties']['name_eng']
    })

df_munic = pd.DataFrame(record_list)

运行后df_munic.shape[0]就会和feature总数完全一致,不会出现记录缺失。

内容的提问来源于stack exchange,提问作者leocanada

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.30 11:27:19