Python 2.7(Windows 10环境)说话人识别技术求助
针对你的语音识别需求的解决方案
Hey there! Since you're on Windows 10 using Python 2.7 and the speech_recognition library, and you want to build a system that only recognizes your voice and what you say, let's break this down into manageable steps—perfect for a beginner!
Core Idea: Combine Speech-to-Text + Speaker Verification
Regular speech recognition only converts audio to text, so you need an extra layer: speaker verification (checking if the voice matches yours). Here's how to do this with tools compatible with Python 2.7:
1. Recommended Libraries for Speaker Verification
These libraries work with Python 2.7 and can handle voice identification:
- pyAudioAnalysis: A versatile library that includes speaker recognition features. It can extract unique voice features from your samples and compare them to incoming audio.
- Install it via pip2:
pip2 install pyaudioanalysis - You'll also need
pyaudiofor recording:pip2 install pyaudio
- Install it via pip2:
- CMU PocketSphinx: While mainly for speech-to-text, it supports speaker identification too. It integrates well with the
speech_recognitionlibrary you're already using.- Install:
pip2 install pocketsphinx
- Install:
2. Step-by-Step Implementation (Using pyAudioAnalysis)
Let's walk through a basic workflow:
- Record Your Voice Samples: Record 5-10 short audio clips (10-30 seconds each) of you speaking in different tones/backgrounds, save them as WAV files in a folder (e.g.,
my_voice_samples). - Train a Speaker Model: Run this code once to create your unique voice template:
from pyAudioAnalysis import audioTrainTest as aT # Train an SVM model with your samples aT.featureAndTrain(["my_voice_samples"], 1.0, 1.0, 0.05, 0.02, "svm", "my_speaker_model", False) - Real-Time Recognition + Verification: Combine this with
speech_recognitionto listen, verify, and transcribe only your voice:import speech_recognition as sr import os from pyAudioAnalysis import audioTrainTest as aT # Initialize the speech recognizer recognizer = sr.Recognizer() with sr.Microphone() as source: print("Speak now...") # Capture audio from the mic audio = recognizer.listen(source) # Save temporary WAV file for speaker verification temp_audio_path = "temp_voice.wav" with open(temp_audio_path, "wb") as f: f.write(audio.get_wav_data()) # Check if the voice matches yours _, prediction_results, _ = aT.fileClassification(temp_audio_path, "my_speaker_model", "svm") # Set a threshold (adjust based on your tests) confidence_threshold = 0.8 if prediction_results[0] > confidence_threshold: print("Recognized your voice! Transcribing...") try: # Convert audio to text (using Google's API) text = recognizer.recognize_google(audio, language="zh-CN") print(f"You said: {text}") except sr.UnknownValueError: print("Sorry, I couldn't understand what you said.") except sr.RequestError as e: print(f"Couldn't connect to the speech service: {e}") else: print("This isn't your voice—ignoring input.") # Clean up temporary file os.remove(temp_audio_path)
3. Pro Tips for Better Accuracy
- Record Samples in Different Scenarios: Include clips with slight background noise, different volumes, and different phrases to make your model more robust.
- Adjust the Confidence Threshold: If the system is too strict (rejects your voice often), lower the threshold a bit; if it accepts others too easily, raise it.
- Note on Python 2.7: Python 2.7 is no longer supported, so if you can, consider upgrading to Python 3.x later—you'll get access to more modern speaker recognition libraries like Resemblyzer or SpeechBrain.
内容的提问来源于stack exchange,提问作者Joel_Developer
相关产品推荐
相关产品推荐

