You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于OpenAI的Node.js说话人分离技术实现咨询

基于OpenAI的Node.js说话人分离技术实现咨询

Hey there! I totally get your frustration—Whisper’s awesome for transcription, but speaker diarization (figuring out which speaker said what) isn’t baked into the OpenAI API directly. Let’s walk through practical ways to pull this off in Node.js, since you’re already handling chunked files with FFmpeg.

核心思路:Whisper + 说话人识别工具搭配使用

Whisper excels at turning audio into text with timestamps, but it can’t identify speakers on its own. The solution is to pair it with tools that specialize in speaker separation, either via cloud APIs (easier for beginners) or local models (more control).


方案1:用云API一站式搞定转录+说话人分离

If you want to keep things simple, services like Deepgram or AssemblyAI offer APIs that handle both transcription and diarization out of the box. Here’s a quick example with Deepgram’s Node.js SDK:

First, install the SDK:

npm install @deepgram/sdk

Then write the code:

const { Deepgram } = require('@deepgram/sdk');
const fs = require('fs');

// Initialize with your API key
const deepgram = new Deepgram('YOUR_DEEPGRAM_API_KEY');

async function transcribeWithSpeakerLabels() {
  // Load your audio file (works with chunked files too, or merge them first for better diarization)
  const audioStream = fs.createReadStream('your-audio-chunk.wav');

  const options = {
    diarize: true, // Enable speaker separation
    model: 'nova-2', // Best model for accuracy
    punctuate: true
  };

  const response = await deepgram.transcription.preRecorded({ stream: audioStream }, options);

  // Extract utterances with speaker labels
  const utterances = response.results.utterances;
  utterances.forEach(utterance => {
    console.log(`Speaker ${utterance.speaker}: ${utterance.transcript}`);
  });
}

transcribeWithSpeakerLabels();

This will give you transcriptions directly tagged with speaker numbers, no extra steps needed.


方案2:结合OpenAI Whisper + 本地说话人识别(更灵活)

If you want to stick with Whisper for transcription and handle diarization separately, here’s how to do it:

  1. Get timestamped transcriptions from Whisper
    First, fetch the transcription with segment timestamps using the OpenAI API:

    const { OpenAI } = require('openai');
    const fs = require('fs');
    const openai = new OpenAI({ apiKey: 'YOUR_OPENAI_KEY' });
    
    async function getWhisperSegments() {
      const transcription = await openai.audio.transcriptions.create({
        file: fs.createReadStream('your-audio-chunk.wav'),
        model: 'whisper-1',
        response_format: 'verbose_json' // This gives us start/end times for each segment
      });
      return transcription.segments;
    }
    
  2. Split audio into segments with FFmpeg
    Use FFmpeg to cut your audio file into individual segments based on Whisper’s timestamps. You can run FFmpeg commands directly from Node.js using child_process:

    const { exec } = require('child_process');
    
    function cutAudioSegment(inputPath, outputPath, startTime, duration) {
      return new Promise((resolve, reject) => {
        exec(
          `ffmpeg -i ${inputPath} -ss ${startTime} -t ${duration} -ac 1 -ar 16000 ${outputPath} -y`,
          (error) => error ? reject(error) : resolve()
        );
      });
    }
    
  3. Identify speakers for each segment
    For local speaker recognition, you can use libraries like tfjs-node to load a pre-trained speaker embedding model (e.g., from SpeechBrain) and cluster segments by speaker. This is more complex, but gives you full control offline.

    A simpler alternative is to use a lightweight npm package like speaker-verification, but note that local models might not be as accurate as cloud APIs.


Pro Tips

  • Merge chunked files first: If your audio is split into chunks, merging them before diarization will help the tool track speakers across chunks more accurately.
  • Specify speaker count: Most diarization tools let you set the number of expected speakers (e.g., diarize_config: { speakers: 2 } in Deepgram), which boosts accuracy.
  • Use WAV format: Convert your audio to WAV (16kHz, mono) before processing—most speaker recognition tools work best with this format.

备注:内容来源于stack exchange,提问作者Zeenath

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.22 11:24:53