You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python 2.7(Windows 10环境)说话人识别技术求助

针对你的语音识别需求的解决方案

Hey there! Since you're on Windows 10 using Python 2.7 and the speech_recognition library, and you want to build a system that only recognizes your voice and what you say, let's break this down into manageable steps—perfect for a beginner!

Core Idea: Combine Speech-to-Text + Speaker Verification

Regular speech recognition only converts audio to text, so you need an extra layer: speaker verification (checking if the voice matches yours). Here's how to do this with tools compatible with Python 2.7:

These libraries work with Python 2.7 and can handle voice identification:

  • pyAudioAnalysis: A versatile library that includes speaker recognition features. It can extract unique voice features from your samples and compare them to incoming audio.
    • Install it via pip2: pip2 install pyaudioanalysis
    • You'll also need pyaudio for recording: pip2 install pyaudio
  • CMU PocketSphinx: While mainly for speech-to-text, it supports speaker identification too. It integrates well with the speech_recognition library you're already using.
    • Install: pip2 install pocketsphinx

2. Step-by-Step Implementation (Using pyAudioAnalysis)

Let's walk through a basic workflow:

  1. Record Your Voice Samples: Record 5-10 short audio clips (10-30 seconds each) of you speaking in different tones/backgrounds, save them as WAV files in a folder (e.g., my_voice_samples).
  2. Train a Speaker Model: Run this code once to create your unique voice template:
    from pyAudioAnalysis import audioTrainTest as aT
    # Train an SVM model with your samples
    aT.featureAndTrain(["my_voice_samples"], 1.0, 1.0, 0.05, 0.02, "svm", "my_speaker_model", False)
    
  3. Real-Time Recognition + Verification: Combine this with speech_recognition to listen, verify, and transcribe only your voice:
    import speech_recognition as sr
    import os
    from pyAudioAnalysis import audioTrainTest as aT
    
    # Initialize the speech recognizer
    recognizer = sr.Recognizer()
    
    with sr.Microphone() as source:
        print("Speak now...")
        # Capture audio from the mic
        audio = recognizer.listen(source)
    
        # Save temporary WAV file for speaker verification
        temp_audio_path = "temp_voice.wav"
        with open(temp_audio_path, "wb") as f:
            f.write(audio.get_wav_data())
    
        # Check if the voice matches yours
        _, prediction_results, _ = aT.fileClassification(temp_audio_path, "my_speaker_model", "svm")
        # Set a threshold (adjust based on your tests)
        confidence_threshold = 0.8
        if prediction_results[0] > confidence_threshold:
            print("Recognized your voice! Transcribing...")
            try:
                # Convert audio to text (using Google's API)
                text = recognizer.recognize_google(audio, language="zh-CN")
                print(f"You said: {text}")
            except sr.UnknownValueError:
                print("Sorry, I couldn't understand what you said.")
            except sr.RequestError as e:
                print(f"Couldn't connect to the speech service: {e}")
        else:
            print("This isn't your voice—ignoring input.")
    
        # Clean up temporary file
        os.remove(temp_audio_path)
    

3. Pro Tips for Better Accuracy

  • Record Samples in Different Scenarios: Include clips with slight background noise, different volumes, and different phrases to make your model more robust.
  • Adjust the Confidence Threshold: If the system is too strict (rejects your voice often), lower the threshold a bit; if it accepts others too easily, raise it.
  • Note on Python 2.7: Python 2.7 is no longer supported, so if you can, consider upgrading to Python 3.x later—you'll get access to more modern speaker recognition libraries like Resemblyzer or SpeechBrain.

内容的提问来源于stack exchange,提问作者Joel_Developer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 08:15:37