You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何识别录制MP3文件中的说话人数并提取候选人文本?附代码求助

解决IBM Watson Speech to Text的说话人识别与文本提取问题

要区分面试录音中的面试官和候选人,你需要开启IBM Watson Speech to Text的**说话人分离(Speaker Diarization)**功能,该功能会自动识别不同说话人并标记对应的文本片段。下面是具体的实现方案和代码修改:

核心修改点

  • 在API调用中添加speaker_labels=True参数,启用说话人识别
  • 修改回调函数,解析返回结果中的speaker_labels字段,提取说话人ID、对应文本及时间戳
  • 基于说话人ID筛选出候选人的文本内容

修改后的完整代码

import json
from os.path import join, dirname
from ibm_watson import SpeechToTextV1
from ibm_watson.websocket import RecognizeCallback, AudioSource
from ibm_cloud_sdk_core.authenticators import IAMAuthenticator

# 初始化认证与服务
authenticator = IAMAuthenticator('your-api-key-here')
speech_to_text = SpeechToTextV1(authenticator=authenticator)
speech_to_text.set_service_url('your-service-url-here')

# 存储所有识别结果(含说话人信息)
transcription_results = []

class MyRecognizeCallback(RecognizeCallback):
    def __init__(self):
        RecognizeCallback.__init__(self)

    def on_data(self, data):
        global transcription_results
        transcription_results.append(data)
        # 打印原始数据(可选,用于调试)
        # print(json.dumps(data, indent=2))

    def on_error(self, error):
        print('Error received: {}'.format(error))

    def on_inactivity_timeout(self, error):
        print('Inactivity timeout: {}'.format(error))

myRecognizeCallback = MyRecognizeCallback()

# 调用API,启用说话人分离
with open(join(dirname(__file__), './sample_2.mp3'), 'rb') as audio_file:
    audio_source = AudioSource(audio_file)
    speech_to_text.recognize_using_websocket(
        audio=audio_source,
        content_type='audio/mp3',
        recognize_callback=myRecognizeCallback,
        model='en-US_BroadbandModel',
        max_alternatives=1,
        speaker_labels=True  # 关键:启用说话人识别
    )

# 解析结果,提取说话人对应文本
def parse_speaker_transcripts(results):
    speaker_transcripts = {}
    # 遍历所有识别片段
    for result in results:
        if 'results' in result:
            for segment in result['results']:
                if not segment['final']:
                    continue
                # 获取当前片段的文本
                transcript = segment['alternatives'][0]['transcript']
                # 获取对应的说话人标签
                if 'speaker_labels' in segment:
                    for label in segment['speaker_labels']:
                        speaker_id = label['speaker']
                        # 初始化说话人字典
                        if speaker_id not in speaker_transcripts:
                            speaker_transcripts[speaker_id] = []
                        # 添加文本片段(可保留时间戳)
                        speaker_transcripts[speaker_id].append({
                            'text': transcript,
                            'start_time': label['start_time'],
                            'end_time': label['end_time']
                        })
    return speaker_transcripts

# 处理并输出结果
speaker_data = parse_speaker_transcripts(transcription_results)

# 统计说话人数
print(f"识别到的说话人数:{len(speaker_data)}")

# 输出每个说话人的完整文本
for speaker_id, transcripts in speaker_data.items():
    full_text = ' '.join([t['text'] for t in transcripts])
    print(f"\n说话人 {speaker_id} 的文本:")
    print(full_text)

# 提取候选人文本(假设候选人是说话人1,可根据实际录音调整)
candidate_speaker_id = 1
if candidate_speaker_id in speaker_data:
    candidate_text = ' '.join([t['text'] for t in speaker_data[candidate_speaker_id]])
    print(f"\n候选人(说话人{candidate_speaker_id})的文本:")
    print(candidate_text)

关键说明

  1. 说话人识别参数:speaker_labels=True是核心,开启后API会返回每个文本片段对应的说话人ID、开始/结束时间戳
  2. 结果解析:speaker_labels字段与results中的文本片段一一对应,需要将两者关联起来
  3. 候选人判定:面试场景中通常可以通过开场对话(比如面试官先提问)判断说话人ID,或者录制前约定测试语句来区分
  4. 模型选择:en-US_BroadbandModel支持说话人分离,若录音是电话音质,可改用en-US_NarrowbandModel

内容的提问来源于stack exchange,提问作者dhruv ranpura

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 21:33:23