WebRTC双端RTP包合并生成对话音频失败问题求助
Fixing Merged WebRTC RTP Streams Only Playing One Client's Audio
Let's break down why your merged audio is only playing one client's voice, and how to fix it:
Root Causes
Your approach has two critical issues that are breaking the merged audio:
- RTP Timestamps Aren't Globally Synced
Each WebRTC peer generates its own RTP timestamp based on its local audio sampling clock. These timestamps don't share a common starting point—so sorting byPacket.Timestampwill misorder packets between the two streams, leading to one stream's audio being placed in the wrong time window (and effectively muted). - Unrecognized SSRCs
Every WebRTC audio stream has a unique SSRC (Synchronization Source Identifier). Youroggwriteris likely designed to handle a single SSRC by default; when it encounters packets from a second SSRC, it either discards them or interprets them as corrupted data from the first stream.
Step-by-Step Solution
1. Sort by Actual Receive Time (Not RTP Timestamp)
Since you're capturing packets in real time, tag each packet with its system receive time when you collect it. This gives you a global, unified timeline to sort all packets correctly.
import "time" // Wrap RTP packets with their capture time type TimedRTPPacket struct { Packet *rtp.Packet RecvTime time.Time } // Collect packets with timestamps var timedList1 []TimedRTPPacket for _, p := range list1 { timedList1 = append(timedList1, TimedRTPPacket{Packet: p, RecvTime: time.Now()}) } var timedList2 []TimedRTPPacket for _, p := range list2 { timedList2 = append(timedList2, TimedRTPPacket{Packet: p, RecvTime: time.Now()}) } // Merge and sort by real capture time timedList := append(timedList1, timedList2...) sort.Slice(timedList, func(i, j int) bool { return timedList[i].RecvTime.Before(timedList[j].RecvTime) }) // Extract sorted RTP packets sortedPackets := make([]*rtp.Packet, len(timedList)) for i, tp := range timedList { sortedPackets[i] = tp.Packet }
2. Decode and Mix Audio from Different SSRCs
Directly writing mixed RTP packets won't work—you need to decode each SSRC's stream separately, then mix the audio into a single stereo (or mono) stream before writing to OGG. Here's how to do it with Opus (the standard WebRTC audio codec):
import ( "github.com/pion/opus" "github.com/pion/oggwriter" ) func mergeAudioStreams(sortedPackets []*rtp.Packet) error { sampleRate := 48000 inputChannels := 1 // WebRTC typically sends mono Opus outputChannels := 2 // Stereo output for separate client audio // Initialize OGG writer for stereo output oggFile, err := oggwriter.New("merged_conversation.ogg", sampleRate, outputChannels) if err != nil { return err } defer oggFile.Close() // Track decoders and PCM buffers for each SSRC decoders := make(map[uint32]*opus.Decoder) pcmBuffers := make(map[uint32][]int16) for _, pkt := range sortedPackets { // Initialize decoder for new SSRCs dec, exists := decoders[pkt.SSRC] if !exists { newDec, err := opus.NewDecoder(sampleRate, inputChannels) if err != nil { return err } decoders[pkt.SSRC] = newDec dec = newDec } // Decode RTP payload to PCM pcmFrame := make([]int16, 1024) // Max Opus frame size for 48kHz decodedSamples, err := dec.Decode(pkt.Payload, pcmFrame, true) // Enable PLC for packet loss recovery if err != nil { // Skip corrupted packets instead of crashing continue } pcmFrame = pcmFrame[:decodedSamples] // Add decoded PCM to the SSRC's buffer pcmBuffers[pkt.SSRC] = append(pcmBuffers[pkt.SSRC], pcmFrame...) // Mix when both SSRCs have PCM data ready if len(pcmBuffers) == 2 { // Get the two SSRC buffers ssrcs := make([]uint32, 0, 2) for s := range pcmBuffers { ssrcs = append(ssrcs, s) } ssrcA, ssrcB := ssrcs[0], ssrcs[1] bufA, bufB := pcmBuffers[ssrcA], pcmBuffers[ssrcB] // Mix up to the length of the shorter buffer mixLength := min(len(bufA), len(bufB)) if mixLength == 0 { continue } // Create stereo PCM: left channel = SSRC A, right channel = SSRC B stereoPCM := make([]int16, mixLength*2) for i := 0; i < mixLength; i++ { stereoPCM[2*i] = bufA[i] // Left channel stereoPCM[2*i+1] = bufB[i] // Right channel } // Write mixed stereo audio to OGG if err := oggFile.Write(stereoPCM); err != nil { return err } // Trim buffers to remove mixed data pcmBuffers[ssrcA] = pcmBuffers[ssrcA][mixLength:] pcmBuffers[ssrcB] = pcmBuffers[ssrcB][mixLength:] } } return nil } func min(a, b int) int { if a < b { return a } return b }
Key Notes
- If you prefer mono output instead of stereo, mix the two PCM streams by averaging samples (e.g.,
mixed[i] = (bufA[i] + bufB[i]) / 2) instead of creating separate channels. - Ensure both streams use the same codec (Opus 48kHz is standard for WebRTC) to avoid decoding mismatches.
- The
PLC(Packet Loss Concealment) flag in the Opus decoder helps smooth over dropped packets, which is common in WebRTC.
内容的提问来源于stack exchange,提问作者Shahriar
相关产品推荐
相关产品推荐

