基于Python的说话人识别工具问询:如何用已有数据识别录音说话人
Python说话人身份识别工具与实现
当然存在适用于Python的类库,能基于已有语音数据完成说话人身份识别,下面是几个实用选项和贴近你需求的代码实现:
常用类库
- pyannote.audio:基于深度学习的语音处理工具,自带预训练模型,可快速实现说话人匹配
- SpeechBrain:开源语音AI框架,提供模块化的说话人验证/识别组件,灵活性强
- librosa + scikit-learn:手动提取语音特征(如梅尔频谱)+ 传统机器学习分类,适合自定义实现
实现示例(基于pyannote.audio)
这个方案贴近你的示例逻辑,用预训练模型快速完成识别:
首先安装依赖:
pip install pyannote.audio sounddevice soundfile scipy
然后编写代码:
import sounddevice as sd import soundfile as sf from pyannote.audio import Pipeline from scipy.spatial.distance import cosine # 替换为你在Hugging Face获取的免费访问令牌 ACCESS_TOKEN = "你的Hugging Face令牌" # 初始化说话人识别管道 pipeline = Pipeline.from_pretrained( "pyannote/speaker-diarization-3.1", use_auth_token=ACCESS_TOKEN ) # 预存的目标说话人语音文件 target_voices = ['obama.wav', 'trump.wav', 'biden.wav'] # 预加载并提取每个说话人的特征嵌入 speaker_embeddings = {} for voice_path in target_voices: # 处理单说话人音频,提取嵌入 diarization = pipeline(voice_path) speaker_segment = next(diarization.itersegments()) speaker_embeddings[voice_path.split('.')[0]] = speaker_segment[2] def record_audio(output_file='recorded.wav', duration=5, samplerate=16000): """录制语音并保存为WAV文件""" print("录音中...") audio_data = sd.rec(int(duration * samplerate), samplerate=samplerate, channels=1) sd.wait() sf.write(output_file, audio_data, samplerate) return output_file # 录制待识别语音 recorded_file = record_audio() # 提取录制语音的说话人嵌入 recorded_diarization = pipeline(recorded_file) recorded_embedding = next(recorded_diarization.itersegments())[2] # 匹配最相似的说话人 min_dist = float('inf') matched_speaker = None for speaker, embedding in speaker_embeddings.items(): dist = cosine(recorded_embedding, embedding) if dist < min_dist: min_dist = dist matched_speaker = speaker print(f"The speaker is likely {matched_speaker}")
轻量实现(基于librosa + scikit-learn)
如果不想依赖深度学习模型,可以用传统机器学习方案:
安装依赖:
pip install librosa scikit-learn sounddevice soundfile
代码实现:
import librosa import numpy as np import sounddevice as sd import soundfile as sf from sklearn.svm import SVC from sklearn.preprocessing import StandardScaler def extract_mfcc_features(file_path): """提取语音的MFCC特征(均值)""" y, sr = librosa.load(file_path, sr=16000) mfcc_features = librosa.feature.mfcc(y=y, sr=sr, n_mfcc=13) return np.mean(mfcc_features, axis=1) def record_audio(output_file='recorded.wav', duration=5, samplerate=16000): print("录音中...") audio_data = sd.rec(int(duration * samplerate), samplerate=samplerate, channels=1) sd.wait() sf.write(output_file, audio_data, samplerate) return output_file # 预存目标语音 target_voices = ['obama.wav', 'trump.wav', 'biden.wav'] # 提取特征并准备训练数据 X = [] labels = [] for voice_path in target_voices: feat = extract_mfcc_features(voice_path) X.append(feat) labels.append(voice_path.split('.')[0]) # 训练分类器 scaler = StandardScaler() X_scaled = scaler.fit_transform(X) classifier = SVC(kernel='linear') classifier.fit(X_scaled, labels) # 录制并识别 recorded_file = record_audio() recorded_feat = extract_mfcc_features(recorded_file) recorded_feat_scaled = scaler.transform([recorded_feat]) matched_speaker = classifier.predict(recorded_feat_scaled)[0] print(f"The speaker is likely {matched_speaker}")
内容的提问来源于stack exchange,提问作者plshelpmeout
相关产品推荐
相关产品推荐

