Python处理超1GB大GeoJSON文件保留重命名指定字段并保持原有结构的方法
方案1:内存足够场景(适合内存≥4G的设备)
直接加载全量JSON修改后回写,逻辑简单易维护:
import json input_path = "data_file_name.json" output_path = "processed.geojson" # 原字段名:重命名后的字段名,不需要重命名则键值保持一致 field_map = { "id": "parcel_id", "area_value": "area_square_meter" } with open(input_path, "r", encoding="utf-8") as f: data = json.load(f) # 仅修改properties字段,其余顶层结构、feature的type/geometry完全保留 for feature in data["features"]: orig_props = feature["properties"] new_props = {new_name: orig_props[old_name] for old_name, new_name in field_map.items()} feature["properties"] = new_props with open(output_path, "w", encoding="utf-8") as f: json.dump(data, f, ensure_ascii=False)
方案2:低内存大文件场景(适合1GB以上超大数据,内存占用≤100MB)
用ijson流式迭代解析+写入,不需要加载全量文件到内存:
import ijson import json input_path = "data_file_name.json" output_path = "processed.geojson" field_map = { "id": "parcel_id", "area_value": "area_square_meter" } with open(input_path, "rb") as in_f, open(output_path, "w", encoding="utf-8") as out_f: # 先读取并写入顶层固定结构 top_type = next(ijson.items(in_f, "type")) crs = next(ijson.items(in_f, "crs")) out_f.write('{"type": "%s", "crs": %s, "features": [' % (top_type, json.dumps(crs, ensure_ascii=False))) # 迭代处理每个feature,处理完直接写入 first_flag = True for feature in ijson.items(in_f, "features.item"): if not first_flag: out_f.write(",") first_flag = False # 仅替换properties字段 orig_props = feature["properties"] new_props = {new_name: orig_props[old_name] for old_name, new_name in field_map.items()} feature["properties"] = new_props out_f.write(json.dumps(feature, ensure_ascii=False)) # 补全JSON结尾结构 out_f.write("]}")
两种方案都完全保留了原始GeoJSON的规范结构,下游程序可以直接识别使用。
内容的提问来源于stack exchange,提问作者Go_ Dmt
相关产品推荐
相关产品推荐

