如何将ConnectyCube P2P视频通话置为保持状态?解决与SpeechRecognizer的音频流冲突问题
Hey there! Let's work through this audio conflict between your ConnectyCube P2P video call and Android's SpeechRecognizer. The core issue here is that both components need exclusive access to the device's audio input stream, which is why you're hitting the ERROR_AUDIO (3) when trying to run SpeechRecognizer mid-call. Let's go through actionable solutions:
1. Simulate Call Hold (Release Audio Resources Properly)
ConnectyCube's SDK doesn't have a built-in "hold" feature, but you can mimic it by properly disabling and re-negotiating your audio track with the peer. Just disabling the track alone isn't enough—you need to update the RTCPeerConnection to notify the other party and fully release the audio capture resource:
currentCall?.let { activeCall -> // Step 1: Disable the local audio track activeCall.mediaStreamManager.localAudioTrack?.setEnabled(false) // Step 2: Remove the audio track from the peer connection activeCall.peerConnection?.removeTrack(activeCall.mediaStreamManager.localAudioTrack?.track) // Step 3: Re-negotiate the session to inform the remote peer activeCall.createOffer { sessionDesc -> activeCall.setLocalDescription(sessionDesc) { activeCall.sendSessionUpdate(sessionDesc) } } }
Once you're done with SpeechRecognizer, reverse these steps to resume audio:
currentCall?.let { activeCall -> activeCall.mediaStreamManager.localAudioTrack?.setEnabled(true) activeCall.peerConnection?.addTrack(activeCall.mediaStreamManager.localAudioTrack?.track, activeCall.mediaStreamManager.localMediaStream) activeCall.createOffer { sessionDesc -> activeCall.setLocalDescription(sessionDesc) { activeCall.sendSessionUpdate(sessionDesc) } } }
This should fully release the audio input stream for SpeechRecognizer to use.
2. Use Cloud-Based Speech Recognition Instead of System SpeechRecognizer
If you don't want to interrupt the call's audio flow, skip the system's SpeechRecognizer entirely. Instead, capture the local audio stream directly from ConnectyCube and send it to a cloud speech-to-text API:
val localAudioTrack = currentCall?.mediaStreamManager?.localAudioTrack localAudioTrack?.track?.addSink(object : AudioSink { override fun onData(buffer: ByteBuffer, timestamp: Long) { // Convert the raw PCM buffer to the format your cloud API requires (e.g., FLAC, WAV) // Send the audio chunk to your cloud speech recognition endpoint } })
This way, you're reusing the same audio stream for both the call and speech recognition, avoiding the exclusive access conflict altogether.
3. Optimize Audio Focus Management
Android's audio focus system can help you coordinate access between the call and SpeechRecognizer. Use AudioManager to request temporary focus before starting SpeechRecognizer, then pause the call's audio:
val audioManager = getSystemService(Context.AUDIO_SERVICE) as AudioManager val focusRequest = AudioFocusRequest.Builder(AudioManager.AUDIOFOCUS_GAIN_TRANSIENT) .setAudioAttributes(AudioAttributes.Builder() .setUsage(AudioAttributes.USAGE_VOICE_COMMUNICATION) .setContentType(AudioAttributes.CONTENT_TYPE_SPEECH) .build()) .setOnAudioFocusChangeListener { focusChange -> if (focusChange == AudioManager.AUDIOFOCUS_LOSS) { // Resume call audio when focus is lost currentCall?.mediaStreamManager?.localAudioTrack?.setEnabled(true) } } .build() val focusResult = audioManager.requestAudioFocus(focusRequest) if (focusResult == AudioManager.AUDIOFOCUS_REQUEST_GRANTED) { // Pause call audio and start SpeechRecognizer currentCall?.mediaStreamManager?.localAudioTrack?.setEnabled(false) speechRecognizer.startListening(yourRecognitionIntent) }
This ensures SpeechRecognizer only gets access to the audio stream when it has the proper focus, and the call resumes automatically once done.
内容的提问来源于stack exchange,提问作者João Fernandes

