You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何缓存分页获取大量数据的NASA REST API调用结果?

针对NASA近地天体API的分页结果缓存方案

1. 本地文件缓存(小型项目快速落地首选)

直接将每页API响应数据保存为本地JSON文件,用页码作为文件名标识(如neo_page_1.json、neo_page_2.json)。调用前先检查对应页码的文件是否存在,存在则读取本地数据,不存在再调用API并写入缓存。

示例代码(Python):

import json
import requests

API_BASE_URL = "https://api.nasa.gov/neo/rest/v1/neo/browse"
API_KEY = "DEMO_KEY"

def get_neo_page(page_num):
    cache_file = f"neo_page_{page_num}.json"
    # 优先读取本地缓存
    try:
        with open(cache_file, 'r') as f:
            return json.load(f)
    except FileNotFoundError:
        # 缓存不存在时调用API
        params = {"api_key": API_KEY, "page": page_num}
        response = requests.get(API_BASE_URL, params=params)
        response.raise_for_status()
        data = response.json()
        # 写入缓存文件
        with open(cache_file, 'w') as f:
            json.dump(data, f, indent=2)
        return data

优点:实现零额外依赖,缓存持久化(重启项目不丢失);缺点:文件数量随页码增多而杂乱,适合数据量较小的场景。

2. 内存缓存(频繁访问场景最优)

用内存字典或轻量缓存库(如functools.lru_cache、cachetools)存储已获取的页码数据,程序运行期间可快速读取,重启后缓存清空。

示例1:用functools.lru_cache实现无过期缓存

import requests
from functools import lru_cache

API_BASE_URL = "https://api.nasa.gov/neo/rest/v1/neo/browse"
API_KEY = "DEMO_KEY"

@lru_cache(maxsize=None)  # maxsize=None表示缓存所有页码
def get_neo_page(page_num):
    params = {"api_key": API_KEY, "page": page_num}
    response = requests.get(API_BASE_URL, params=params)
    response.raise_for_status()
    return response.json()

示例2:用cachetools实现带过期时间的缓存

import requests
from cachetools import TTLCache

# 最多缓存100页,每30天自动过期(近地天体数据更新频率极低)
cache = TTLCache(maxsize=100, ttl=30*24*3600)

def get_neo_page(page_num):
    if page_num in cache:
        return cache[page_num]
    params = {"api_key": API_KEY, "page": page_num}
    response = requests.get(API_BASE_URL, params=params)
    response.raise_for_status()
    data = response.json()
    cache[page_num] = data
    return data

优点:读取速度极快,适合高频查询;缺点:缓存不持久,仅适合临时缓存或配合持久化方案使用。

3. 轻量数据库缓存(结构化持久化需求)

用SQLite这类轻量数据库创建缓存表,存储页码、JSON数据、缓存时间,查询时优先读数据库,无缓存则调用API并写入。

示例代码(Python):

import sqlite3
import json
import requests

API_BASE_URL = "https://api.nasa.gov/neo/rest/v1/neo/browse"
API_KEY = "DEMO_KEY"

# 初始化缓存数据库
def init_cache_db():
    conn = sqlite3.connect('neo_cache.db')
    cursor = conn.cursor()
    cursor.execute('''
        CREATE TABLE IF NOT EXISTS neo_pages (
            page_num INTEGER PRIMARY KEY,
            data TEXT,
            cached_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP
        )
    ''')
    conn.commit()
    conn.close()

init_cache_db()

def get_neo_page(page_num):
    # 查询数据库缓存
    conn = sqlite3.connect('neo_cache.db')
    cursor = conn.cursor()
    cursor.execute('SELECT data FROM neo_pages WHERE page_num = ?', (page_num,))
    result = cursor.fetchone()
    if result:
        conn.close()
        return json.loads(result[0])
    # 调用API获取数据
    params = {"api_key": API_KEY, "page": page_num}
    response = requests.get(API_BASE_URL, params=params)
    response.raise_for_status()
    data = response.json()
    # 写入数据库缓存
    cursor.execute('''
        INSERT INTO neo_pages (page_num, data)
        VALUES (?, ?)
    ''', (page_num, json.dumps(data)))
    conn.commit()
    conn.close()
    return data

优点:结构化管理缓存,支持持久化,可通过cached_at字段实现过期清理;缺点:需要编写数据库操作代码,适合数据量稍大、需长期保存缓存的场景。

实用缓存策略补充

  • 预缓存:先调用一次API获取total总条数,计算出总页码后,在项目启动时批量调用所有页码并缓存,后续直接使用本地数据。
  • 缓存过期:由于近地天体数据更新频率低,可设置30天以上的过期时间,定期清理过期缓存或在调用时检查缓存时间,超期则重新获取。
  • 错误过滤:API调用失败时(如状态码非200),不要将错误结果写入缓存,避免后续读取无效数据。

内容的提问来源于stack exchange,提问作者ViDanMaster

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 04:23:24