You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用http.get获取URL对应PDF Buffer的代码优化方案咨询

批量获取PDF Buffer实现优化建议

Node.js 原生 http 模块没有支持传入URL数组的批量请求接口,Promise.all 本身已经是并发IO场景下的标准实现,不会额外增加耗时,你可以从以下几个方向优化现有逻辑,提升性能和稳定性:


1. 增加并发控制,避免触发限流或连接溢出

如果后续URL数量超过10个,全并发请求会占用大量TCP连接,很容易触发目标站点的限流策略,也可能导致本地端口耗尽。可以通过简单的并发数限制控制请求节奏,示例实现如下:

// 简单实现并发控制器,也可以直接用成熟的p-limit库
class ConcurrencyLimit {
  private queue: (() => Promise<any>)[] = []
  private active = 0
  constructor(private limit: number) {}

  add<T>(fn: () => Promise<T>): Promise<T> {
    return new Promise((resolve, reject) => {
      this.queue.push(() => fn().then(resolve).catch(reject))
      this.run()
    })
  }

  private run() {
    while (this.active < this.limit && this.queue.length) {
      const task = this.queue.shift()!
      this.active++
      task().finally(() => {
        this.active--
        this.run()
      })
    }
  }
}

// 使用示例,限制最多同时3个请求
const limit = new ConcurrencyLimit(3)
const BuffersPdfs = await Promise.all(
  urlPdf.map(url => limit.add(() => this.getBufferOfPDF(url)))
)

2. 完善getBufferOfPDF方法的鲁棒性

现有实现存在3个明显缺陷:

  • 仅支持http协议的URL,遇到https链接会报错
  • 没有请求超时机制,异常场景下会永久挂住Promise
  • 没有错误重试机制,单次网络波动就会导致整个流程失败

优化后的实现参考:

import * as http from 'http'
import * as https from 'https'
import { URL } from 'url'

private static async getBufferOfPDF(url: string, retryCount = 3, timeout = 10000): Promise<Buffer> {
  return new Promise((resolve, reject) => {
    // 自动适配http/https协议
    const urlObj = new URL(url)
    const client = urlObj.protocol === 'https:' ? https : http

    const req = client.get(url, (res) => {
      if (res.statusCode < 200 || res.statusCode >= 300) {
        return reject(new Error('statusCode=' + res.statusCode));
      }
      const chunkData: Buffer[] = []
      res.on('data', chunk => chunkData.push(chunk))
      res.on('end', () => resolve(Buffer.concat(chunkData)))
      res.on('error', reject)
    })

    // 超时处理
    req.setTimeout(timeout, () => {
      req.destroy()
      reject(new Error('Request timeout'))
    })

    req.on('error', e => {
      console.error("Error:", e)
      reject(e.message);
    })
    req.end()
  }).catch(err => {
    // 重试逻辑
    if (retryCount > 0) {
      console.log(`请求${url}失败,剩余重试次数${retryCount}`)
      return this.getBufferOfPDF(url, retryCount - 1, timeout)
    }
    throw err
  })
}

3. 内存占用优化

如果单PDF体积大、数量多,全量Buffer存在内存中会导致内存占用过高,可以调整流程为边下载边合并,用流的方式处理PDF,不需要把所有文件都加载到内存,合并完成后直接将流上传到S3,能大幅降低大文件场景下的内存消耗。


4. 错误降级可选优化

如果业务允许个别PDF拉取失败时不中断整个流程,可以调整Promise.all为Promise.allSettled,执行完成后过滤出成功的Buffer再做合并,避免单个请求失败导致所有任务作废。


内容的提问来源于stack exchange,提问作者Programmer89

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.07 13:48:01