如何使用Firestore写入第三方API新增天气数据并忽略已存在的旧数据
问题根因
ArrayUnion的去重逻辑是严格匹配完全相等的元素,天气数据对象如果有字段顺序、空格、浮点精度等微小差异,就会被判定为新元素写入,无法达到预期去重效果。- 逐元素循环发起读写请求,单条数据就要消耗1读1写,数据量累积后成本必然上涨。
- 文档ID包含动态的
current time字段,会导致同一个城市同一天的数据被写入不同ID的文档,无法找到已有数据完成追加,反而出现重复写入甚至覆盖的问题。
实现方案
1. 优化存储结构
- 固定文档ID规则为
城市+日期格式,例如Weather-Austin, Texas-2021-08-25,移除ID中动态的current time字段,保证同一城市同一天的所有数据都存入固定ID的文档。 - 文档内部弃用数组结构,改用时间字符串为键的嵌套字典存储天气数据,天生自带键去重能力,不需要额外判断重复:
{ "weather": { "9:30 am PST": {"weather": "95 degrees F"}, "11:30 am PST": {"weather": "72 degrees F"} } }
2. 批量处理降低读写次数
先按文档ID分组所有待写入数据,每个文档仅发起1次读请求,判断有新增数据再发起写请求,大幅降低调用量:
from collections import defaultdict import firebase_admin from firebase_admin import firestore # 全局初始化一次Firestore客户端,不要在循环内重复创建 if not firebase_admin._apps: firebase_admin.initialize_app() db = firestore.client() # 按文档ID分组所有待写入数据 grouped_data = defaultdict(dict) for wd in write_data: collection_name = wd[0] # 从原始标识中提取城市+日期生成固定文档ID city_part = wd[1].split(" from 8 am")[0].strip() date_part = wd[1].split("Date: ")[1].strip() fixed_doc_id = f"{city_part}-{date_part}" time_key = wd[2].pop("time") grouped_data[(collection_name, fixed_doc_id)][time_key] = wd[2] # 逐个文档处理 for (coll_name, doc_id), time_map in grouped_data.items(): ref = db.collection(coll_name).document(doc_id) doc = ref.get() if doc.exists: existing_weather = doc.to_dict().get("weather", {}) # 筛选出未存储的新增时间点数据 new_entries = {k: v for k, v in time_map.items() if k not in existing_weather} if new_entries: # 批量更新嵌套字段,仅写入新增数据 ref.update({f"weather.{k}": v for k, v in new_entries.items()}) else: ref.set({"weather": time_map})
3. 额外成本优化
- 本地缓存当日已写入的城市+时间点标识,程序运行过程中无需重复读取Firestore判断重复,每天零点清空缓存即可,可几乎消除所有读请求。
- 每次请求API后记录最新的时间点,下次请求仅处理晚于该时间点的新数据,省去全量去重步骤。
内容的提问来源于stack exchange,提问作者Tim Gorer
相关产品推荐
相关产品推荐

