You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure Speech Service(v3.2)批量转录speaker diarization质量问题及模型咨询

Azure Speech Service v3.2批量转录问题及技术问询解答

当前问题

  • 使用Azure Speech Service v3.2进行批量转录
  • Speaker diarization(说话人分离)无法区分男女说话人
  • 此前功能正常
  • 使用单声道音频文件

当前配置

{
    "properties": {
        "diarizationEnabled": true,
        "wordLevelTimestampsEnabled": true,
        "displayFormWordLevelTimestampsEnabled": true,
        "channels": [0, 1],
        "diarization": {
            "speakers": {
                "minCount": 1,
                "maxCount": 25
            }
        }
    },
    "locale": "de-DE"
}

当前模型

使用基础模型:acf3c487-5d8c-4a4a-8241-f508cb5f2059(德国中西部德语)

当前代码

import sys
import requests
import time
import ast
import os
import copy
import boto3
from datetime import datetime, timedelta
import swagger_client
from azure.storage.blob import BlobClient, generate_blob_sas, BlobSasPermissions


def transcribe_from_single_blob(self, uri, job_id, language, properties):
    """
    Transcribe a single audio file located at `uri` using the settings specified in `properties`
    using the base model for the specified locale.
    """

    transcription_definition = swagger_client.Transcription(
        display_name=str(job_id),
        description='Transciption with Azure Base Model',
        locale=language,
        content_urls=[uri],
        properties=properties
    ) 

def transcribe(self, blob_uri, job_id, language):
    logging.info("Starting transcription client...")

    # configure API key authorization: subscription_key
    configuration = swagger_client.Configuration()
    configuration.api_key["Ocp-Apim-Subscription-Key"] = self.SUBSCRIPTION_KEY
    configuration.host = f"https://{self.SERVICE_REGION}.api.cognitive.microsoft.com/speechtotext/v3.2"

    properties = swagger_client.TranscriptionProperties()
    properties.profanity_filter_mode = "None"  
    properties.time_to_live = "PT48H"  
    properties.punctuation_mode = "Automatic"
    properties.word_level_timestamps_enabled = True
    properties.display_form_word_level_timestamps_enabled = True


    if self.number_speaker != 1:
        properties.diarization_enabled = True
        properties.diarization = swagger_client.DiarizationProperties(swagger_client.DiarizationSpeakersProperties(min_count=1, max_count=(25 if self.number_speaker == 0 else self.number_speaker)))
    else:
        properties.diarization_enabled = False

    # create the client object and authenticate
    client = swagger_client.ApiClient(configuration)

    # create an instance of the transcription api class
    api =swagger_client.CustomSpeechTranscriptionsApi(api_client=client)
    transcription_definition = self.transcribe_from_single_blob(blob_uri, job_id, language, properties)
    created_transcription, status, headers = api.transcriptions_create_with_http_info(transcription=transcription_definition) 

技术问询解答

1. 如何提升speaker diarization的识别质量?

  • 修正声道配置:当前配置里channels设为[0,1],但实际用的是单声道音频,这会干扰服务的说话人分离逻辑,改成[0]匹配单声道输入。
  • 缩小说话人数量范围:当前minCount=1、maxCount=25的范围太宽泛,服务需要额外计算来匹配大量可能的说话人。如果已知场景是男女两位说话人,直接把minCount和maxCount都设为2,能大幅提升识别精准度。
  • 检查音频质量:确保音频无明显背景噪音、回声,说话人语音清晰、音量稳定,低质量音频会严重影响说话人特征提取。
  • 核对模型版本:确认当前使用的基础模型是否为最新版本,旧模型可能存在说话人分离的性能缺陷。

2. 如何为德语批量转录场景选择正确的基础模型?

  • 匹配地域方言:当前使用的德国中西部德语模型,适合对应方言的音频;如果是奥地利德语、瑞士德语等其他方言,需要选择对应地域的模型。
  • 优先批量转录优化模型:Azure针对批量转录场景提供了专门优化的基础模型,这类模型在长时间音频处理、说话人分离上表现更优,可在Azure门户Speech服务的模型列表中筛选带「批量转录」标签的德语模型。
  • 测试对比:选取典型音频样本,用不同德语模型测试转录和说话人分离效果,选择准确率最高的模型。

3. 常规基础模型与批量转录专用模型之间有什么差异?

  • 优化方向不同:常规基础模型侧重实时转录场景,核心是低延迟;批量转录专用模型针对长时间非实时音频处理优化,重点提升转录准确率、说话人分离精度,以及大音量音频队列的处理稳定性。
  • 资源占用与速度:批量模型处理速度相对慢,但能调用更多资源处理复杂音频;常规模型资源占用低、响应快,适合实时交互场景。
  • 功能支持:批量模型通常支持更多高级功能,比如长音频精准说话人分离、多轮对话上下文识别优化;常规模型在这类复杂场景的表现不如批量模型。
  • 适用场景:常规模型适合实时语音转写、语音助手等场景;批量模型适合录音文件归档、会议转录、媒体内容批量处理等非实时场景。

内容的提问来源于stack exchange,提问作者JustusFer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.15 22:29:59