Android通话中Speech Recognizer语音转文字及双向语音交互方案求助
Hey there! Let’s break down how to build this real-time call audio workflow—capturing speech from your speaker, converting it to text, processing that text, then turning it back to speech and sending it through your mic during an active call. Here’s a practical, actionable breakdown:
First, let’s map out the end-to-end pipeline you’ll need:
- Capture audio from your speaker (the other caller’s voice)
- Convert that audio stream to text in real time (STT)
- Run your custom text processing logic (cleanup, summarization, translation, etc.)
- Convert the processed text back to speech (TTS)
- Inject that speech audio into your system’s microphone input so the other caller hears it
1. Capture Speaker Audio
To record the audio coming out of your speaker (not just your mic), you’ll need a virtual loopback audio device or platform-specific audio capture APIs:
- Windows: Use
WasapiLoopbackCapture(via libraries likepyaudioorsounddevice) or tools like VB-Cable to create a loopback device. - macOS: Use BlackHole or Soundflower to route system audio to a virtual input.
- Linux: Use PulseAudio’s loopback module (
pactl load-module module-loopback).
In Python, you can use sounddevice to list and select the loopback device:
import sounddevice as sd # List all audio devices to find your loopback/speaker capture device print(sd.query_devices()) # Set the device index (replace with your loopback device's index) loopback_device_idx = 2
2. Real-Time Speech-to-Text (STT)
For low-latency, offline STT (critical for call scenarios), OpenAI’s Whisper is a great choice—it’s fast, accurate, and runs locally. Use the smaller models (base or small) for real-time performance:
import whisper import sounddevice as sd import sys # Load Whisper model (base is fast enough for real time) model = whisper.load_model("base") sample_rate = 16000 # Whisper expects 16kHz audio def audio_callback(indata, frames, time, status): if status: print(f"Audio status error: {status}", file=sys.stderr) # Convert the audio chunk to text transcription = model.transcribe(indata.flatten(), fp16=False) raw_text = transcription["text"].strip() if raw_text: # Pass to your text processing function processed_text = process_text(raw_text) # Send processed text to TTS and mic tts_and_inject_mic(processed_text) # Start capturing from the loopback device with sd.InputStream(device=loopback_device_idx, samplerate=sample_rate, callback=audio_callback): input("Press Enter to stop the workflow...\n")
3. Custom Text Processing
Write your own process_text function to handle whatever logic you need—cleanup, filtering, summarization, translation, etc. Example:
def process_text(raw_text): # Example: Clean up transcription and add a custom response prefix cleaned_text = raw_text.replace("um", "").strip() return f"[Processed] {cleaned_text}"
4. Text-to-Speech (TTS) & Inject to Microphone
For offline TTS, Coqui TTS is a solid open-source option. You’ll need to route the TTS output to your virtual microphone device so the call software picks it up:
from TTS.api import TTS import sounddevice as sd # Load a TTS model (English example) tts = TTS(model_name="tts_models/en/ljspeech/tacotron2-DDC") mic_device_idx = 3 # Replace with your virtual mic device index def tts_and_inject_mic(text): # Generate speech audio from text audio = tts.tts(text) # Play the audio directly to the virtual mic device sd.play(audio, samplerate=22050, device=mic_device_idx) sd.wait() # Wait for playback to finish before next chunk
- Optimize for Latency: Use smaller STT/TTS models, reduce audio chunk sizes, and run processing in separate threads to avoid delays that disrupt the call flow.
- Virtual Device Setup: Make sure your call software (Zoom, Teams, etc.) is set to use the virtual microphone as its input device, and your regular speaker as output (so you can hear the caller too).
- Platform Testing: Audio APIs vary across Windows/macOS/Linux—test your device indices and capture/injection logic thoroughly on your target platform.
- Error Handling: Add checks for empty transcriptions, audio glitches, and model loading failures to keep the workflow stable during calls.
If you run into specific issues—like trouble detecting loopback devices or latency spikes—feel free to share more details, and we can troubleshoot further!
内容的提问来源于stack exchange,提问作者Malik Basit

