WebRTC音频流保存与字节数组获取实现方案咨询
Hey there! Let's break down how to capture that WebRTC audio stream into a byte array or save it as a playable audio file. You're already halfway there with your existing AudioFrameReady handler—here's how to extend it:
1. Get a Byte Array from the AudioFrame
The audioData IntPtr points to raw PCM audio data. To convert this into a managed byte array, we'll use Marshal.Copy (from System.Runtime.InteropServices). You'll need to calculate the total number of bytes based on the frame's properties:
track.AudioFrameReady += (AudioFrame frame) => { IntPtr audioDataPtr = frame.audioData; // Calculate byte count per sample and total bytes for the frame int bytesPerSample = frame.bitsPerSample / 8; int totalFrameBytes = frame.samplesPerChannel * frame.channels * bytesPerSample; // Initialize a byte array and copy the unmanaged data into it byte[] audioFrameBytes = new byte[totalFrameBytes]; System.Runtime.InteropServices.Marshal.Copy(audioDataPtr, audioFrameBytes, 0, totalFrameBytes); // Now you can use audioFrameBytes for processing, caching, or real-time handling };
2. Save the Stream as a Playable WAV File
WebRTC sends audio in raw PCM format, which isn't directly playable. To save it as a WAV file, we need to accumulate PCM frames over time and add a WAV file header (which describes the audio format for media players).
First, add class-level variables to track audio properties and accumulate data:
private MemoryStream _pcmAccumulator; private int _audioSampleRate; private int _audioChannels; private int _audioBitsPerSample;
Then, initialize these when the remote track is added:
pc.AudioTrackAdded += (RemoteAudioTrack track) => { var readerBuffer = track.CreateReadBuffer(); // Capture core audio properties (adjust property names if your API uses different ones) _audioSampleRate = track.SampleRate; // WebRTC commonly uses 48000 Hz _audioChannels = track.Channels; // Typically 1 (mono) or 2 (stereo) _audioBitsPerSample = 16; // 16-bit PCM is standard for WebRTC _pcmAccumulator = new MemoryStream(); track.AudioFrameReady += (AudioFrame frame) => { // Reuse the byte array logic from step 1 IntPtr audioDataPtr = frame.audioData; int bytesPerSample = frame.bitsPerSample / 8; int totalFrameBytes = frame.samplesPerChannel * frame.channels * bytesPerSample; byte[] audioFrameBytes = new byte[totalFrameBytes]; System.Runtime.InteropServices.Marshal.Copy(audioDataPtr, audioFrameBytes, 0, totalFrameBytes); // Add the current frame to our accumulated PCM data _pcmAccumulator.Write(audioFrameBytes, 0, audioFrameBytes.Length); }; };
Finally, add a method to convert the accumulated PCM data to a WAV file:
private void SaveAudioToWav(string outputFilePath) { if (_pcmAccumulator == null || _pcmAccumulator.Length == 0) { Console.WriteLine("No audio data available to save!"); return; } long totalPcmBytes = _pcmAccumulator.Length; long totalWavFileSize = totalPcmBytes + 44; // WAV header takes 44 bytes long byteRate = _audioSampleRate * _audioChannels * (_audioBitsPerSample / 8); using (var fileStream = new FileStream(outputFilePath, FileMode.Create)) using (var writer = new BinaryWriter(fileStream)) { // Write the WAV file header writer.Write(new[] { 'R', 'I', 'F', 'F' }); writer.Write(totalWavFileSize); writer.Write(new[] { 'W', 'A', 'V', 'E' }); writer.Write(new[] { 'f', 'm', 't', ' ' }); writer.Write(16); // PCM format subchunk size writer.Write((short)1); // Audio format: PCM = 1 writer.Write((short)_audioChannels); writer.Write(_audioSampleRate); writer.Write(byteRate); writer.Write((short)(_audioChannels * (_audioBitsPerSample / 8))); // Block alignment writer.Write((short)_audioBitsPerSample); writer.Write(new[] { 'd', 'a', 't', 'a' }); writer.Write(totalPcmBytes); // Write the accumulated PCM data _pcmAccumulator.Position = 0; _pcmAccumulator.CopyTo(fileStream); } // Clean up resources to avoid memory leaks _pcmAccumulator.Dispose(); _pcmAccumulator = null; }
Quick Tips
- Verify frame properties: Double-check that
bitsPerSample,samplesPerChannel, andchannelsmatch the actual data from your WebRTC implementation (16-bit, 48000 Hz, mono is most common). - Real-time handling: If you need to stream audio to a file as it arrives instead of saving all at once, you can write the byte arrays directly to a file stream (just make sure to write the WAV header first before adding PCM data).
- Resource cleanup: Always dispose of streams when you're done to free up memory.
内容的提问来源于stack exchange,提问作者AmirHossein

