You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用jq流式处理大型GeoJSON并转换为CSV格式?

问题描述

我有一些大型GeoJSON文件(比如美国建筑足迹数据集里的文件),需要转成.csv格式方便用Postgres的\copy快速导入,文件通过STDIN传入。这些GeoJSON里的几何类型都是Polygon,我只需要提取每个多边形的第一个点(经纬度),不需要完整几何和属性信息。

用jq处理小文件(<1.4GB)没问题,命令如下:

jq '.features | map(.geometry.coordinates) | map(.[]) | map(first) | .[] | {"long": first, "lat": last} | [.long, .lat] | @csv' small.geojson

但处理超大文件时会因为内存不足被系统Kill。尝试用--stream参数,但要么用法不对,要么速度极慢(跑3小时还没结束)。试过这个命令:

cat sample.geojson | jq --stream "fromstream(1|truncate_stream(inputs))" | jq ' map(.geometry.coordinates) | map(.[]) | map(first) | .[] | {"long": first, "lat": last} | [.long, .lat] | @csv'

样本文件能跑,但大文件不行,求解决办法。

样本GeoJSON

{
  "type": "FeatureCollection",
  "features": [
    {
      "type": "Feature",
      "geometry": {
        "type": "Polygon",
        "coordinates": [
          [
            [
              -84.959634,
              32.421887
            ],
            [
              -84.95982,
              32.421889
            ],
            [
              -84.959822,
              32.421797
            ],
            [
              -84.959767,
              32.421796
            ],
            [
              -84.959767,
              32.421771
            ],
            [
              -84.959636,
              32.421769
            ],
            [
              -84.959634,
              32.421887
            ]
          ]
        ]
      },
      "properties": {
        "release": 2,
        "capture_dates_range": "3/26/2020-7/22/2020"
      }
    },
    {
      "type": "Feature",
      "geometry": {
        "type": "Polygon",
        "coordinates": [
          [
            [
              -84.959636,
              32.42095
            ],
            [
              -84.959715,
              32.42095
            ],
            [
              -84.959714,
              32.420984
            ],
            [
              -84.959816,
              32.420985
            ],
            [
              -84.959818,
              32.420849
            ],
            [
              -84.959637,
              32.420848
            ],
            [
              -84.959636,
              32.42095
            ]
          ]
        ]
      },
      "properties": {
        "release": 2,
        "capture_dates_range": "3/26/2020-7/22/2020"
      }
    },
    {
      "type": "Feature",
      "geometry": {
        "type": "Polygon",
        "coordinates": [
          [
            [
              -84.959998,
              32.235231
            ],
            [
              -84.959877,
              32.235231
            ],
            [
              -84.959877,
              32.235288
            ],
            [
              -84.959998,
              32.235288
            ],
            [
              -84.959998,
              32.235231
            ]
          ]
        ]
      },
      "properties": {
        "release": 1,
        "capture_dates_range": ""
      }
    },
    {
      "type": "Feature",
      "geometry": {
        "type": "Polygon",
        "coordinates": [
          [
            [
              -84.960253,
              32.422248
            ],
            [
              -84.960069,
              32.422245
            ],
            [
              -84.960067,
              32.422321
            ],
            [
              -84.960165,
              32.422323
            ],
            [
              -84.960164,
              32.422364
            ],
            [
              -84.96025,
              32.422365
            ],
            [
              -84.960253,
              32.422248
            ]
          ]
        ]
      },
      "properties": {
        "release": 2,
        "capture_dates_range": "3/26/2020-7/22/2020"
      }
    },
    {
      "type": "Feature",
      "geometry": {
        "type": "Polygon",
        "coordinates": [
          [
            [
              -84.961602,
              32.419206
            ],
            [
              -84.961599,
              32.419354
            ],
            [
              -84.961707,
              32.419355
            ],
            [
              -84.961708,
              32.419291
            ],
            [
              -84.961794,
              32.419292
            ],
            [
              -84.961796,
              32.419208
            ],
            [
              -84.961602,
              32.419206
            ]
          ]
        ]
      },
      "properties": {
        "release": 2,
        "capture_dates_range": "3/26/2020-7/22/2020"
      }
    }
  ]
}
解决方案

核心问题是你之前的--stream用法仍在重建整个FeatureCollection,导致内存占用过高。正确的做法是用jq流式处理直接定位目标坐标点,无需加载完整JSON结构。

高效流式处理命令

用这条单jq命令处理,全程流式读取,内存占用极低:

jq --stream -n '
  def get_first_point:
    foreach inputs as $item (
      {};
      if $item[0][0] == "features" and ($item[0][2] | tonumber?) != null and $item[0][3] == "geometry" and $item[0][4] == "coordinates" and $item[0][5] == 0 and $item[0][6] == 0 then
        $item[1] as $coord
        | {long: $coord[0], lat: $coord[1]}
        | [.long, .lat]
        | @csv
        | print
        | .
      else
        .
      end
    );
  get_first_point
' large.geojson

命令说明

  1. --stream -n:-n让jq从空输入启动,配合--stream逐块读取JSON流,避免一次性加载整个文件
  2. 路径匹配逻辑:精准定位到每个Feature的geometry.coordinates[0][0](即Polygon第一个环的第一个点),匹配到后立即提取经纬度转CSV输出
  3. 无内存累积:处理完一个点就输出,不会缓存整个数据集,内存占用始终维持在低水平

样本验证

用你的sample.geojson测试,会输出:

-84.959634,32.421887
-84.959636,32.42095
-84.959998,32.235231
-84.960253,32.422248
-84.961602,32.419206

和原小文件命令输出一致,但可轻松处理GB级大文件。

替代方案(极端大文件场景)

如果jq速度仍不满足需求,可使用Python的ijson库做流式处理,内存占用同样极低:

import ijson
import sys
import csv

writer = csv.writer(sys.stdout)
for feature in ijson.items(sys.stdin, 'features.item'):
    coord = feature['geometry']['coordinates'][0][0]
    writer.writerow([coord[0], coord[1]])

运行方式:

python3 extract_points.py < large.geojson > output.csv

内容的提问来源于stack exchange,提问作者defuneste

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 09:02:02