You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

WebRTC双端RTP包合并生成对话音频失败问题求助

Fixing Merged WebRTC RTP Streams Only Playing One Client's Audio

Let's break down why your merged audio is only playing one client's voice, and how to fix it:

Root Causes

Your approach has two critical issues that are breaking the merged audio:

  1. RTP Timestamps Aren't Globally Synced
    Each WebRTC peer generates its own RTP timestamp based on its local audio sampling clock. These timestamps don't share a common starting point—so sorting by Packet.Timestamp will misorder packets between the two streams, leading to one stream's audio being placed in the wrong time window (and effectively muted).
  2. Unrecognized SSRCs
    Every WebRTC audio stream has a unique SSRC (Synchronization Source Identifier). Your oggwriter is likely designed to handle a single SSRC by default; when it encounters packets from a second SSRC, it either discards them or interprets them as corrupted data from the first stream.

Step-by-Step Solution

1. Sort by Actual Receive Time (Not RTP Timestamp)

Since you're capturing packets in real time, tag each packet with its system receive time when you collect it. This gives you a global, unified timeline to sort all packets correctly.

import "time"

// Wrap RTP packets with their capture time
type TimedRTPPacket struct {
    Packet   *rtp.Packet
    RecvTime time.Time
}

// Collect packets with timestamps
var timedList1 []TimedRTPPacket
for _, p := range list1 {
    timedList1 = append(timedList1, TimedRTPPacket{Packet: p, RecvTime: time.Now()})
}

var timedList2 []TimedRTPPacket
for _, p := range list2 {
    timedList2 = append(timedList2, TimedRTPPacket{Packet: p, RecvTime: time.Now()})
}

// Merge and sort by real capture time
timedList := append(timedList1, timedList2...)
sort.Slice(timedList, func(i, j int) bool {
    return timedList[i].RecvTime.Before(timedList[j].RecvTime)
})

// Extract sorted RTP packets
sortedPackets := make([]*rtp.Packet, len(timedList))
for i, tp := range timedList {
    sortedPackets[i] = tp.Packet
}

2. Decode and Mix Audio from Different SSRCs

Directly writing mixed RTP packets won't work—you need to decode each SSRC's stream separately, then mix the audio into a single stereo (or mono) stream before writing to OGG. Here's how to do it with Opus (the standard WebRTC audio codec):

import (
    "github.com/pion/opus"
    "github.com/pion/oggwriter"
)

func mergeAudioStreams(sortedPackets []*rtp.Packet) error {
    sampleRate := 48000
    inputChannels := 1 // WebRTC typically sends mono Opus
    outputChannels := 2 // Stereo output for separate client audio

    // Initialize OGG writer for stereo output
    oggFile, err := oggwriter.New("merged_conversation.ogg", sampleRate, outputChannels)
    if err != nil {
        return err
    }
    defer oggFile.Close()

    // Track decoders and PCM buffers for each SSRC
    decoders := make(map[uint32]*opus.Decoder)
    pcmBuffers := make(map[uint32][]int16)

    for _, pkt := range sortedPackets {
        // Initialize decoder for new SSRCs
        dec, exists := decoders[pkt.SSRC]
        if !exists {
            newDec, err := opus.NewDecoder(sampleRate, inputChannels)
            if err != nil {
                return err
            }
            decoders[pkt.SSRC] = newDec
            dec = newDec
        }

        // Decode RTP payload to PCM
        pcmFrame := make([]int16, 1024) // Max Opus frame size for 48kHz
        decodedSamples, err := dec.Decode(pkt.Payload, pcmFrame, true) // Enable PLC for packet loss recovery
        if err != nil {
            // Skip corrupted packets instead of crashing
            continue
        }
        pcmFrame = pcmFrame[:decodedSamples]

        // Add decoded PCM to the SSRC's buffer
        pcmBuffers[pkt.SSRC] = append(pcmBuffers[pkt.SSRC], pcmFrame...)

        // Mix when both SSRCs have PCM data ready
        if len(pcmBuffers) == 2 {
            // Get the two SSRC buffers
            ssrcs := make([]uint32, 0, 2)
            for s := range pcmBuffers {
                ssrcs = append(ssrcs, s)
            }
            ssrcA, ssrcB := ssrcs[0], ssrcs[1]
            bufA, bufB := pcmBuffers[ssrcA], pcmBuffers[ssrcB]

            // Mix up to the length of the shorter buffer
            mixLength := min(len(bufA), len(bufB))
            if mixLength == 0 {
                continue
            }

            // Create stereo PCM: left channel = SSRC A, right channel = SSRC B
            stereoPCM := make([]int16, mixLength*2)
            for i := 0; i < mixLength; i++ {
                stereoPCM[2*i] = bufA[i]     // Left channel
                stereoPCM[2*i+1] = bufB[i]   // Right channel
            }

            // Write mixed stereo audio to OGG
            if err := oggFile.Write(stereoPCM); err != nil {
                return err
            }

            // Trim buffers to remove mixed data
            pcmBuffers[ssrcA] = pcmBuffers[ssrcA][mixLength:]
            pcmBuffers[ssrcB] = pcmBuffers[ssrcB][mixLength:]
        }
    }
    return nil
}

func min(a, b int) int {
    if a < b {
        return a
    }
    return b
}

Key Notes

  • If you prefer mono output instead of stereo, mix the two PCM streams by averaging samples (e.g., mixed[i] = (bufA[i] + bufB[i]) / 2) instead of creating separate channels.
  • Ensure both streams use the same codec (Opus 48kHz is standard for WebRTC) to avoid decoding mismatches.
  • The PLC (Packet Loss Concealment) flag in the Opus decoder helps smooth over dropped packets, which is common in WebRTC.

内容的提问来源于stack exchange,提问作者Shahriar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 13:27:56