如何识别录制MP3文件中的说话人数并提取候选人文本?附代码求助
解决IBM Watson Speech to Text的说话人识别与文本提取问题
要区分面试录音中的面试官和候选人,你需要开启IBM Watson Speech to Text的**说话人分离(Speaker Diarization)**功能,该功能会自动识别不同说话人并标记对应的文本片段。下面是具体的实现方案和代码修改:
核心修改点
- 在API调用中添加
speaker_labels=True参数,启用说话人识别 - 修改回调函数,解析返回结果中的
speaker_labels字段,提取说话人ID、对应文本及时间戳 - 基于说话人ID筛选出候选人的文本内容
修改后的完整代码
import json from os.path import join, dirname from ibm_watson import SpeechToTextV1 from ibm_watson.websocket import RecognizeCallback, AudioSource from ibm_cloud_sdk_core.authenticators import IAMAuthenticator # 初始化认证与服务 authenticator = IAMAuthenticator('your-api-key-here') speech_to_text = SpeechToTextV1(authenticator=authenticator) speech_to_text.set_service_url('your-service-url-here') # 存储所有识别结果(含说话人信息) transcription_results = [] class MyRecognizeCallback(RecognizeCallback): def __init__(self): RecognizeCallback.__init__(self) def on_data(self, data): global transcription_results transcription_results.append(data) # 打印原始数据(可选,用于调试) # print(json.dumps(data, indent=2)) def on_error(self, error): print('Error received: {}'.format(error)) def on_inactivity_timeout(self, error): print('Inactivity timeout: {}'.format(error)) myRecognizeCallback = MyRecognizeCallback() # 调用API,启用说话人分离 with open(join(dirname(__file__), './sample_2.mp3'), 'rb') as audio_file: audio_source = AudioSource(audio_file) speech_to_text.recognize_using_websocket( audio=audio_source, content_type='audio/mp3', recognize_callback=myRecognizeCallback, model='en-US_BroadbandModel', max_alternatives=1, speaker_labels=True # 关键:启用说话人识别 ) # 解析结果,提取说话人对应文本 def parse_speaker_transcripts(results): speaker_transcripts = {} # 遍历所有识别片段 for result in results: if 'results' in result: for segment in result['results']: if not segment['final']: continue # 获取当前片段的文本 transcript = segment['alternatives'][0]['transcript'] # 获取对应的说话人标签 if 'speaker_labels' in segment: for label in segment['speaker_labels']: speaker_id = label['speaker'] # 初始化说话人字典 if speaker_id not in speaker_transcripts: speaker_transcripts[speaker_id] = [] # 添加文本片段(可保留时间戳) speaker_transcripts[speaker_id].append({ 'text': transcript, 'start_time': label['start_time'], 'end_time': label['end_time'] }) return speaker_transcripts # 处理并输出结果 speaker_data = parse_speaker_transcripts(transcription_results) # 统计说话人数 print(f"识别到的说话人数:{len(speaker_data)}") # 输出每个说话人的完整文本 for speaker_id, transcripts in speaker_data.items(): full_text = ' '.join([t['text'] for t in transcripts]) print(f"\n说话人 {speaker_id} 的文本:") print(full_text) # 提取候选人文本(假设候选人是说话人1,可根据实际录音调整) candidate_speaker_id = 1 if candidate_speaker_id in speaker_data: candidate_text = ' '.join([t['text'] for t in speaker_data[candidate_speaker_id]]) print(f"\n候选人(说话人{candidate_speaker_id})的文本:") print(candidate_text)
关键说明
- 说话人识别参数:
speaker_labels=True是核心,开启后API会返回每个文本片段对应的说话人ID、开始/结束时间戳 - 结果解析:
speaker_labels字段与results中的文本片段一一对应,需要将两者关联起来 - 候选人判定:面试场景中通常可以通过开场对话(比如面试官先提问)判断说话人ID,或者录制前约定测试语句来区分
- 模型选择:
en-US_BroadbandModel支持说话人分离,若录音是电话音质,可改用en-US_NarrowbandModel
内容的提问来源于stack exchange,提问作者dhruv ranpura
相关产品推荐
相关产品推荐

