Python解析GeoJSON转DataFrame时出现数据行数缺失问题
问题描述
我正尝试从市级行政区域GeoJSON文件中提取韩国各自治市的韩文、英文名称,所用Python代码如下:
import json import pandas as pd Korean_municipalities = json.load(open('skorea-municipalities-2018-geo.json', 'r')) munic_map_eng = {} for feature in Korean_municipalities['features']: feature['id'] = feature['properties']['name_eng'] munic_map_eng[feature['properties']['name']] = feature['id'] df_munic = pd.DataFrame(list(munic_map_eng.items()))
运行后出现数据缺失问题:
- 统计得
len(Korean_municipalities['features']) = 250,即文件内共包含250个市级行政单元,但最终生成的DataFrame对象df_munic维度为df_munic.shape = (227,2),共缺失23条记录。 - 将相同逻辑代码用于省级、sub-municipality级行政区域名称提取时:
- sub-municipality级共3504个行政单元,最终生成的DataFrame仅3142行,存在同类缺失问题
- 省级共17个行政单元,提取结果完全正常
问题原因
核心原因是Python字典的键唯一性规则:
你用韩文行政区名feature['properties']['name']作为字典的键,韩国省级行政区共17个,本身没有重名,所以不会出现键覆盖,结果完全正常。但市级、sub-municipality级行政区存在大量跨上级区域的重名情况,遍历feature时,后出现的重名条目会直接覆盖字典中已存在的同键值,最终字典条目数必然小于实际feature总数,转成DataFrame就会出现记录缺失。
修正方案
不要用字典做中间存储触发去重逻辑,直接遍历所有feature时将需要的字段存入列表,再转成DataFrame即可避免覆盖问题,参考代码:
import json import pandas as pd Korean_municipalities = json.load(open('skorea-municipalities-2018-geo.json', 'r')) record_list = [] for feature in Korean_municipalities['features']: # 按需提取字段,可加入行政编码、上级行政区名等字段做唯一区分 record_list.append({ "name_korean": feature['properties']['name'], "name_english": feature['properties']['name_eng'] }) df_munic = pd.DataFrame(record_list)
运行后df_munic.shape[0]就会和feature总数完全一致,不会出现记录缺失。
内容的提问来源于stack exchange,提问作者leocanada
相关产品推荐
相关产品推荐

