You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Google Cloud Speech-to-Text API邮箱地址转写优化咨询

Fixing Email Address Transcription in Google Cloud Speech-to-Text

Great question! I've run into this exact issue when working with Google Cloud Speech-to-Text for transcribing structured identifiers like email addresses. The model defaults to spelling out "at" and "dot" because it's optimized for conversational speech, but there are several ways to guide it to output the actual email format instead. Here are the most effective methods:

  • Use Speech Context Phrase Hints
    This is the quickest win. By providing the model with specific phrases or patterns you expect to hear, you tell it to prioritize those exact formats over spelled-out alternatives. For example, if you're transcribing habjan@domain.com, add that exact string (or broader patterns like *@domain.com for multiple emails under the same domain) to the speech_contexts parameter in your API request.

    Here's a Python client library snippet to implement this:

    from google.cloud import speech_v1p1beta1 as speech
    
    client = speech.SpeechClient()
    
    # Define context hints to prioritize email formats
    speech_contexts = [{"phrases": ["habjan@domain.com", "*@domain.com"]}]
    
    config = speech.RecognitionConfig(
        encoding=speech.RecognitionConfig.AudioEncoding.LINEAR16,
        sample_rate_hertz=16000,
        language_code="en-US",
        speech_contexts=speech_contexts,
        enable_automatic_punctuation=True,  # Critical for symbol recognition
    )
    
  • Enable Automatic Punctuation
    Turning on automatic punctuation helps the model recognize that @ and . are part of structured text rather than spoken words. This pairs perfectly with phrase hints to ensure the model uses the correct symbols instead of "at" or "dot". Enable it via the enable_automatic_punctuation flag in your recognition config, as shown in the snippet above.

  • Switch to the Command-and-Search Model
    The default default model is built for general conversation. For short, precise content like emails, commands, or search queries, use the command_and_search model. It's trained to prioritize structured identifiers and will be far more likely to output @ and . instead of their spelled-out equivalents.

    Update your config to include this model:

    config = speech.RecognitionConfig(
        # ... existing settings
        model="command_and_search",
    )
    
  • Leverage Custom Classes (For Scalable Workflows)
    If you need to handle a large volume of emails or multiple consistent patterns, create a custom class. Define a class like email_address and add regex-like patterns that match email structures (e.g., [a-z]+@[a-z]+.[a-z]+). This tells the model to categorize incoming speech as an email and format it correctly automatically.

内容的提问来源于stack exchange,提问作者HABJAN

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 08:47:12