如何缓存分页获取大量数据的NASA REST API调用结果?
针对NASA近地天体API的分页结果缓存方案
1. 本地文件缓存(小型项目快速落地首选)
直接将每页API响应数据保存为本地JSON文件,用页码作为文件名标识(如neo_page_1.json、neo_page_2.json)。调用前先检查对应页码的文件是否存在,存在则读取本地数据,不存在再调用API并写入缓存。
示例代码(Python):
import json import requests API_BASE_URL = "https://api.nasa.gov/neo/rest/v1/neo/browse" API_KEY = "DEMO_KEY" def get_neo_page(page_num): cache_file = f"neo_page_{page_num}.json" # 优先读取本地缓存 try: with open(cache_file, 'r') as f: return json.load(f) except FileNotFoundError: # 缓存不存在时调用API params = {"api_key": API_KEY, "page": page_num} response = requests.get(API_BASE_URL, params=params) response.raise_for_status() data = response.json() # 写入缓存文件 with open(cache_file, 'w') as f: json.dump(data, f, indent=2) return data
优点:实现零额外依赖,缓存持久化(重启项目不丢失);缺点:文件数量随页码增多而杂乱,适合数据量较小的场景。
2. 内存缓存(频繁访问场景最优)
用内存字典或轻量缓存库(如functools.lru_cache、cachetools)存储已获取的页码数据,程序运行期间可快速读取,重启后缓存清空。
示例1:用functools.lru_cache实现无过期缓存
import requests from functools import lru_cache API_BASE_URL = "https://api.nasa.gov/neo/rest/v1/neo/browse" API_KEY = "DEMO_KEY" @lru_cache(maxsize=None) # maxsize=None表示缓存所有页码 def get_neo_page(page_num): params = {"api_key": API_KEY, "page": page_num} response = requests.get(API_BASE_URL, params=params) response.raise_for_status() return response.json()
示例2:用cachetools实现带过期时间的缓存
import requests from cachetools import TTLCache # 最多缓存100页,每30天自动过期(近地天体数据更新频率极低) cache = TTLCache(maxsize=100, ttl=30*24*3600) def get_neo_page(page_num): if page_num in cache: return cache[page_num] params = {"api_key": API_KEY, "page": page_num} response = requests.get(API_BASE_URL, params=params) response.raise_for_status() data = response.json() cache[page_num] = data return data
优点:读取速度极快,适合高频查询;缺点:缓存不持久,仅适合临时缓存或配合持久化方案使用。
3. 轻量数据库缓存(结构化持久化需求)
用SQLite这类轻量数据库创建缓存表,存储页码、JSON数据、缓存时间,查询时优先读数据库,无缓存则调用API并写入。
示例代码(Python):
import sqlite3 import json import requests API_BASE_URL = "https://api.nasa.gov/neo/rest/v1/neo/browse" API_KEY = "DEMO_KEY" # 初始化缓存数据库 def init_cache_db(): conn = sqlite3.connect('neo_cache.db') cursor = conn.cursor() cursor.execute(''' CREATE TABLE IF NOT EXISTS neo_pages ( page_num INTEGER PRIMARY KEY, data TEXT, cached_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP ) ''') conn.commit() conn.close() init_cache_db() def get_neo_page(page_num): # 查询数据库缓存 conn = sqlite3.connect('neo_cache.db') cursor = conn.cursor() cursor.execute('SELECT data FROM neo_pages WHERE page_num = ?', (page_num,)) result = cursor.fetchone() if result: conn.close() return json.loads(result[0]) # 调用API获取数据 params = {"api_key": API_KEY, "page": page_num} response = requests.get(API_BASE_URL, params=params) response.raise_for_status() data = response.json() # 写入数据库缓存 cursor.execute(''' INSERT INTO neo_pages (page_num, data) VALUES (?, ?) ''', (page_num, json.dumps(data))) conn.commit() conn.close() return data
优点:结构化管理缓存,支持持久化,可通过cached_at字段实现过期清理;缺点:需要编写数据库操作代码,适合数据量稍大、需长期保存缓存的场景。
实用缓存策略补充
- 预缓存:先调用一次API获取
total总条数,计算出总页码后,在项目启动时批量调用所有页码并缓存,后续直接使用本地数据。 - 缓存过期:由于近地天体数据更新频率低,可设置30天以上的过期时间,定期清理过期缓存或在调用时检查缓存时间,超期则重新获取。
- 错误过滤:API调用失败时(如状态码非200),不要将错误结果写入缓存,避免后续读取无效数据。
内容的提问来源于stack exchange,提问作者ViDanMaster
相关产品推荐
相关产品推荐

