Node.js高并发场景下fetch拉取资源频繁超时问题及优化咨询
问题背景
使用Node.js v20.16.0编写高并发请求测试代码,运行时频繁出现超时,并发量越高概率越大,换用node-fetch也存在同样问题。测试发现部分特定链接仅在并发请求时超时,单独请求200ms内即可完成(超时设置为5秒)。
测试代码如下:
// NodeJs version: v20.16.0 async function fetchFile(url: string, timeout: number): Promise<Response> { const controller = new AbortController() const id = setTimeout(() => controller.abort('timeout: ' + url), timeout) return fetch(url, {signal: controller.signal}) .finally(() => clearTimeout(id)) } const testUrl = 'https://lf3-cdn-tos.bytecdntp.com/cdn/expire-1-M/KaTeX/0.15.2/contrib/copy-tex.min.js' const count = 100 const timeout = 5000 async function test() { for (let i = 0; i < count; i++) { try { await Promise.all( new Array(count).fill(testUrl).map(it => fetchFile(it, timeout)) ) console.log('success: ' + i) } catch (e: any) { console.error(`error[${i}]: ${e}`) } } } test()
核心疑问
- 为什么高并发请求时会频繁超时,单独请求却正常?
- 如何在最大化并发量的前提下避免这类超时问题?
一、成因分析
1. Node.js默认HTTP连接限制
Node.js的http/https模块默认对同一域名的并发连接数有限制(默认6个),当发起100个并发请求到同一域名时,大部分请求会处于等待队列中,超过5秒就触发了超时逻辑。即使是不同域名,系统的文件描述符、端口资源也有上限,短时间大量创建连接会导致资源耗尽,引发超时。
2. 目标服务器限流/队列溢出
很多CDN或服务器会对来自同一IP的并发请求做限流,当请求量超过阈值时,服务器会直接丢弃请求或让请求排队,排队时间过长就会触发超时。
3. 连接池未高效复用
默认情况下,Node.js的HTTP Agent会复用连接,但如果请求过于密集,连接池来不及创建足够的连接,或者请求结束后连接回收不及时,都会导致新请求需要重新建立TCP连接,耗时增加进而超时。
二、解决方案
1. 调整HTTP Agent的并发参数
通过自定义Agent提高同一域名的并发连接数,同时启用长连接复用:
import https from 'https'; const customAgent = new https.Agent({ keepAlive: true, // 启用长连接复用,减少TCP握手开销 maxSockets: 100, // 同一域名最大并发连接数,按需调整 maxFreeSockets: 20, // 空闲时保持的连接数,避免频繁创建连接 }); // 修改fetchFile函数,传入自定义Agent async function fetchFile(url: string, timeout: number): Promise<Response> { const controller = new AbortController() const id = setTimeout(() => controller.abort('timeout: ' + url), timeout) return fetch(url, { signal: controller.signal, agent: customAgent }) .finally(() => clearTimeout(id)) }
注:maxSockets不要设置过大,避免触发服务器限流或耗尽本地资源。
2. 实现全局请求限流/排队
即使是请求不同域名,也需要控制全局并发量,避免系统资源耗尽。可以用简单的队列控制并发数:
async function limitedConcurrentRequest<T>(tasks: (() => Promise<T>)[], limit: number): Promise<T[]> { const results: T[] = []; const executing: Promise<void>[] = []; for (const task of tasks) { const p = task().then(res => results.push(res)); executing.push(p); if (executing.length >= limit) { const finished = await Promise.race(executing); executing.splice(executing.indexOf(finished), 1); } } await Promise.all(executing); return results; } // 修改test函数使用限流逻辑 async function test() { const tasks = new Array(count).fill(testUrl).map(it => () => fetchFile(it, timeout)); for (let i = 0; i < count; i++) { try { await limitedConcurrentRequest(tasks, 30); // 控制全局并发量为30 console.log('success: ' + i) } catch (e: any) { console.error(`error[${i}]: ${e}`) } } }
3. 优化超时计时逻辑
当前的超时是从请求发起开始计算,但等待连接的时间也会被计入,导致还没真正发送请求就超时。可以调整为请求真正建立连接后才开始计时:
async function fetchFile(url: string, timeout: number): Promise<Response> { const controller = new AbortController() let timeoutId: NodeJS.Timeout; const response = await fetch(url, { signal: controller.signal, agent: customAgent }); // 连接建立后才启动超时计时(针对响应慢的情况) timeoutId = setTimeout(() => controller.abort('response timeout: ' + url), timeout); // 读取响应内容后清除超时 const resClone = response.clone(); await resClone.text(); clearTimeout(timeoutId); return response; }
4. 针对限流域名做请求打散
如果某些域名容易触发限流,可以给请求添加随机延迟,避免请求集中发起:
async function fetchFile(url: string, timeout: number): Promise<Response> { // 添加0-100ms的随机延迟,打散请求时间 await new Promise(resolve => setTimeout(resolve, Math.random() * 100)); const controller = new AbortController() const id = setTimeout(() => controller.abort('timeout: ' + url), timeout) return fetch(url, { signal: controller.signal, agent: customAgent }) .finally(() => clearTimeout(id)) }
5. 监控连接状态排查泄漏
通过Agent的事件监控连接情况,排查是否有连接泄漏:
customAgent.on('free', (socket) => { console.log(`连接已释放,当前空闲连接数:${customAgent.freeSockets.size}`); }); customAgent.on('close', (socket) => { console.log('连接已关闭'); });
内容的提问来源于stack exchange,提问作者kmar

