You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Node.js中向OpenAI Whisper API传入File对象的正确方法

Node.js环境下调用Whisper API转写URL音频文件的问题

需求

我希望从URL下载文件后调用Whisper API进行转写。

官方文档示例

按照OpenAI官方文档的示例编写代码:

const resp = await openai.createTranscription(
  fs.createReadStream("audio.mp3"),
  "whisper-1"
);

我的具体实现

public static async transcribeFromPublicUrl({ url, format }: { url: string; format: string }) {
    const now = new Date().toISOString();
    const filePath = `${this.tmpdir}/${now}.${format}`;
    try {
      const response = await axios.get<Stream>(url, {
        responseType: 'stream',
      });
      const fileStream = fs.createWriteStream(filePath);
      response.data.pipe(fileStream);

      await new Promise((resolve, reject) => {
        fileStream.on('finish', resolve);
        fileStream.on('error', reject);
      });

      const transcriptionResponse = await 
      this.openai.createTranscription(fs.readFileSync(filePath), 'whisper');
      return { success: true, response: transcriptionResponse };
    } catch (error) {
      console.error('Failed to download the file:', error);
      return { success: false, error: error };
    }
  }

遇到的错误

执行代码时出现TypeScript类型错误:

Argument of type 'Buffer' is not assignable to parameter of type 'File'.
  Type 'Buffer' is missing the following properties from type 'File': lastModified, name, webkitRelativePath, size, and 5 more.ts(2345)

尝试的解决方法

尝试将Buffer转换为File对象:

...
 const file = new File([fs.readFileSync(filePath)], now, { type: `audio/${format}` });
 const transcriptionResponse = await this.openai.createTranscription(file, 'whisper');
...

虽然TypeScript不再报错,但Node.js环境不支持浏览器的File API,运行时会报错。

查看库定义后的问题

查看OpenAI库的类型定义,createTranscription方法要求第一个参数为File类型:

/**
     *
     * @summary Transcribes audio into the input language.
     * @param {File} file The audio file to transcribe, in one of these formats: mp3, mp4, mpeg, mpga, m4a, wav, or webm.
     * @param {string} model ID of the model to use. Only `whisper-1` is currently available.
     * @param {string} [prompt] An optional text to guide the model's style or continue a previous audio segment. The prompt should match the audio language.
     * @param {string} [responseFormat] The format of the transcript output, in one of these options: json, text, srt, verbose_json, or vtt.
     * @param {number} [temperature] The sampling temperature, between 0 and 1. Higher values like 0.8 will make the output more random, while lower values like 0.2 will make it more focused and deterministic. If set to 0, the model will use log probability to automatically increase the temperature until certain thresholds are hit.
     * @param {string} [language] The language of the input audio. Supplying the input language in ISO-639-1 format will improve accuracy and latency.
     * @param {*} [options] Override http request option.
     * @throws {RequiredError}
     * @memberof OpenAIApi
*/
createTranscription(file: File, model: string, prompt?: string, responseFormat?: string, temperature?: number, language?: string, options?: AxiosRequestConfig): Promise<import("axios").AxiosResponse<CreateTranscriptionResponse, any>>;

问题总结

Node.js环境中无法使用浏览器的File API,但OpenAI库的createTranscription方法要求传入File类型参数,该如何解决?


内容的提问来源于stack exchange,提问作者Peter Toth

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 10:17:55