You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

从ArcGIS REST服务批量获取并合并JSON数据遇问题求助

批量下载ArcGIS REST服务图形数据并合并JSON文件

我需要从ArcGIS REST服务下载带图形(shapes)的JSON数据,该服务每次最多允许请求1000个图形。我想通过循环每次请求1000条数据,然后将所有请求到的图形合并到单个JSON文件中。目前我先生成第一批1000条数据的JSON,再用后续请求的数据更新它,但遇到了很多问题。以下是我的代码:

import requests
import json

# URL of the ArcGIS Server Map Service Layer
url = "https://map.sitr.regione.sicilia.it/gis/rest/services/catasto/cartografia_catastale/MapServer/4/query"
# Define the file path and name for the new Geopackage
geopackage_path = 'C:/Users/Daniele/OneDrive/PhD/GIS + ESOM'
n_Items = 8237133
maxitem = 999
# Identify the number of requests
# n_requests = int(n_Items / maxitem)
# Only for test purpose ( To eliminate once code is OK)
n_requests = 10
print(f"Total number of request is {n_requests}")

for i in range(n_requests):
    lowbound = i * 1000 + 1
    upbound = lowbound + maxitem
    if (n_Items - upbound) < maxitem:
        upbound = n_Items

    # Parameters for the ArcGIS Server Map Service Layer request
    params = {
        "where": f"objectid BETWEEN {lowbound} AND {upbound}",
        "outFields": "*",
        "returnGeometry": "true",
        "f": "json"
    }

    response = requests.get(url, params=params)

    data = json.loads(response.text)

    if i == 0:
        with open('data1.json', 'w', encoding='utf-8') as f:
            json.dump(data, f, ensure_ascii=False, indent=4)
    else:
        with open('data1.json', 'r+', encoding='utf-8') as f:
            file_data = json.load(f)
            file_data.update(data)
            f.seek(0)
            json.dump(file_data, f, indent=4)

问题分析与修正方案

核心问题

你的代码使用file_data.update(data)来合并数据,这是错误的。ArcGIS REST返回的JSON结构中,图形数据存储在features数组里,而update方法会直接覆盖file_data中的同名字段(比如fields、spatialReference),而不是追加features数组的内容。此外,用r+模式写入时如果新内容长度短于旧内容,会残留旧数据。

修正后的代码

import requests
import json

# ArcGIS服务地址
url = "https://map.sitr.regione.sicilia.it/gis/rest/services/catasto/cartografia_catastale/MapServer/4/query"
total_items = 8237133
batch_size = 1000  # 每次请求1000条,对应where条件的区间长度
output_file = 'data1.json'
# 测试用,实际运行时注释掉下面一行,启用计算逻辑
n_requests = 10
# n_requests = (total_items + batch_size - 1) // batch_size  # 计算总请求次数,向上取整

print(f"总请求次数: {n_requests}")

# 初始化合并后的数据
merged_data = None

for i in range(n_requests):
    # 计算当前批次的objectid区间
    start_id = i * batch_size + 1
    end_id = min(start_id + batch_size - 1, total_items)
    
    params = {
        "where": f"objectid BETWEEN {start_id} AND {end_id}",
        "outFields": "*",
        "returnGeometry": "true",
        "f": "json"
    }

    try:
        response = requests.get(url, params=params, timeout=30)
        response.raise_for_status()  # 检查请求是否成功
        data = response.json()  # 直接用response.json()解析,更简洁
        
        if 'error' in data:
            print(f"请求第{i+1}批次失败: {data['error']['message']}")
            continue

        if merged_data is None:
            # 第一次请求,初始化合并数据,保留元数据和第一批features
            merged_data = data
        else:
            # 后续请求,追加features数组
            merged_data['features'].extend(data['features'])
        
        print(f"已完成第{i+1}批次,累计获取{len(merged_data['features'])}条数据")

    except requests.exceptions.RequestException as e:
        print(f"第{i+1}批次请求出错: {str(e)}")
    except json.JSONDecodeError:
        print(f"第{i+1}批次返回数据不是有效JSON")

# 写入最终合并后的JSON
if merged_data:
    with open(output_file, 'w', encoding='utf-8') as f:
        json.dump(merged_data, f, ensure_ascii=False, indent=4)
    print(f"所有数据已合并并保存到{output_file}")
else:
    print("未获取到任何有效数据")

关键改进点

  • 合并逻辑修正:不再用update,而是直接追加features数组,保留第一次请求的元数据(fields、spatialReference等),这些元数据所有批次都是一致的,无需重复覆盖。
  • 区间计算优化:用min(start_id + batch_size - 1, total_items)直接处理最后一批次的边界,逻辑更清晰。
  • 异常处理:增加请求超时、HTTP错误、JSON解析错误的捕获,避免程序中途崩溃。
  • 文件写入优化:最后一次性写入所有数据,而不是每次循环都修改文件,更高效且避免残留问题。
  • 更简洁的JSON解析:使用response.json()替代json.loads(response.text)。

内容的提问来源于stack exchange,提问作者Daniele Mosso

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 16:02:47