Cloud Run部署Node.js上传Firestore遇DEADLINE_EXCEEDED错误求助
Cloud Run Node.js应用Firestore上传DEADLINE_EXCEEDED超时问题解决
问题概述
部署在Cloud Run的Node.js应用向Firestore上传大数据集时,频繁触发4 DEADLINE_EXCEEDED: Deadline exceeded after 102.893s错误。尽管数据最终能成功写入,但上传耗时极长,且错误出现后后续所有批次都会触发相同超时。
环境配置
- 资源:内存8GIB,CPU 2核
- 请求配置:超时1800秒,单实例最大并发请求80,仅请求处理时分配CPU
- 实例数量:最小0,最大25
核心代码片段
上传逻辑
import { firestore } from "firebase-admin"; import { db } from "../firebase/config"; import { areAllValuesNullOrEmpty } from "./areAllValuesNullOrEmpty"; import { removeNullValues } from "./filterNullValues"; const BATCH_SIZE = 450; const delay = (ms: number) => new Promise((resolve) => setTimeout(resolve, ms)); export const uploadToFirestore = async ( userRefId: string, collectionName: string, docId: string, data: object, retries: number = 3, delayMs: number = 1000 ): Promise<void> => { try { const filteredData = removeNullValues(data); if (areAllValuesNullOrEmpty(filteredData)) { return; } const patientRef = db.collection("patients").doc(userRefId); const summaryRef = patientRef.collection(collectionName).doc(docId); await uploadInBatches(summaryRef, filteredData); console.log(`Data uploaded to ${collectionName}/${docId}`); } catch (error) { console.error(`Failed to upload data to ${collectionName}/${docId}`, error); if (retries > 0) { console.log( `Retrying upload to ${collectionName}/${docId} (${retries} retries left)...` ); await delay(delayMs); await uploadToFirestore( userRefId, collectionName, docId, data, retries - 1, delayMs ); } else { console.error( `Exhausted retries for uploading data to ${collectionName}/${docId}` ); } } }; const uploadInBatches = async ( docRef: firestore.DocumentReference, data: object ): Promise<void> => { const batch = db.batch(); batch.set(docRef, data); await batch.commit(); }; // uploadActivityToFirestore 代码省略
调用逻辑
const handleDaily = async (payload: TerraPayload) => { const data = payload.data as unknown as Daily[]; const user = payload.user as unknown as TerraUser; const userRefId = user.reference_id; for (const dailyData of data) { const startTime = dailyData.metadata.start_time; const endTime = dailyData.metadata.end_time; const collectionId = generateCollectionId(startTime, endTime); const collectionsData = generateDailyCollections(dailyData); for (const [collectionName, collectionData] of Object.entries( collectionsData )) { await uploadToFirestore( userRefId, collectionName, collectionId, collectionData ); } } };
已尝试方案(未解决)
- 使用BulkWriter、BatchWrite、事务
- 设置请求延迟
- 拆分数据为更小块
- 有限次数重试
问题根源分析
- Firestore客户端默认超时过短:Node.js客户端默认gRPC请求超时为60秒,大文档写入或高网络延迟场景下,会触发超时重试,累积后出现100+秒的超时错误。
- Cloud Run CPU策略限制:「仅请求处理时分配CPU」模式下,长时间上传任务可能遭遇CPU节流,导致处理速度变慢,延长写入时间。
- 低效写入逻辑:
uploadInBatches并未实现真正批量写入,每个调用仅处理单个文档,且串行执行所有上传任务,导致请求排队、耗时累积。 - 单文档体积过大:若上传文档接近或超过Firestore 1MB的单文档上限,写入耗时会显著增加,极易触发超时。
具体解决方案
1. 调整Firestore客户端超时配置
初始化客户端时设置更长的超时,匹配Cloud Run的请求超时:
import admin from 'firebase-admin'; admin.initializeApp({ credential: admin.credential.applicationDefault(), }); const db = admin.firestore(); // 设置超时为5分钟(300秒) db.settings({ timeoutSeconds: 300, });
2. 修改Cloud Run CPU分配策略
将CPU配置改为「始终分配CPU」,确保长时间上传任务全程获得稳定的CPU资源,避免节流导致的处理变慢。
3. 重构批量上传逻辑
实现真正的批量写入,减少请求次数:
const uploadInBatches = async ( operations: Array<{ref: firestore.DocumentReference, data: object, type: 'set' | 'update'}> ): Promise<void> => { const batchSize = 450; // 不超过Firestore批量操作上限500 let currentBatch = db.batch(); const commitPromises: Promise<void>[] = []; for (let i = 0; i < operations.length; i++) { const op = operations[i]; op.type === 'set' ? currentBatch.set(op.ref, op.data) : currentBatch.update(op.ref, op.data); if ((i + 1) % batchSize === 0) { commitPromises.push(currentBatch.commit()); currentBatch = db.batch(); } } if (operations.length % batchSize !== 0) { commitPromises.push(currentBatch.commit()); } await Promise.all(commitPromises); };
同时将串行调用改为批量收集后处理:
const handleDaily = async (payload: TerraPayload) => { const data = payload.data as unknown as Daily[]; const user = payload.user as unknown as TerraUser; const userRefId = user.reference_id; const allOperations: Array<{ref: firestore.DocumentReference, data: object, type: 'set'}> = []; for (const dailyData of data) { const startTime = dailyData.metadata.start_time; const endTime = dailyData.metadata.end_time; const collectionId = generateCollectionId(startTime, endTime); const collectionsData = generateDailyCollections(dailyData); for (const [collectionName, collectionData] of Object.entries(collectionsData)) { const filteredData = removeNullValues(collectionData); if (areAllValuesNullOrEmpty(filteredData)) continue; const docRef = db.collection("patients").doc(userRefId) .collection(collectionName).doc(collectionId); allOperations.push({ref: docRef, data: filteredData, type: 'set'}); } } await uploadInBatches(allOperations); };
4. 拆分过大文档
若单文档数据量接近1MB上限,将数据拆分为子集合的多个文档存储,避免单文档写入耗时过长。
5. 对齐Cloud Run与Firestore区域
确保两者部署区域一致,跨区域部署会增加网络延迟,显著提升写入耗时。
6. 优化重试机制
重试时创建全新批次,避免复用错误状态的资源:
export const uploadToFirestore = async ( userRefId: string, collectionName: string, docId: string, data: object, retries: number = 3, delayMs: number = 1000 ): Promise<void> => { try { const filteredData = removeNullValues(data); if (areAllValuesNullOrEmpty(filteredData)) { return; } const docRef = db.collection("patients").doc(userRefId) .collection(collectionName).doc(docId); // 每次尝试创建新批次 const batch = db.batch(); batch.set(docRef, filteredData); await batch.commit(); console.log(`Data uploaded to ${collectionName}/${docId}`); } catch (error) { console.error(`Failed to upload data to ${collectionName}/${docId}`, error); if (retries > 0) { console.log(`Retrying upload to ${collectionName}/${docId} (${retries} retries left)...`); await delay(delayMs); await uploadToFirestore(userRefId, collectionName, docId, data, retries - 1, delayMs); } else { console.error(`Exhausted retries for uploading data to ${collectionName}/${docId}`); } } };
内容的提问来源于stack exchange,提问作者Hamdi Al Masalmeh
相关产品推荐
相关产品推荐

