You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何向Telegram Bot发送语音指令?技术实现问询

实现Telegram Bot语音指令的完整方案

当然可以实现这个需求!我来一步步给你拆解怎么打造一个能接收语音、转文本、翻译后返回的Telegram Bot:

1. 先让Bot能监听语音消息

不管你用哪个Telegram Bot开发框架(比如常用的pyTelegramBotAPI或者python-telegram-bot),第一步都是配置Bot来捕获语音消息事件。这里以pyTelegramBotAPI为例,写个基础的监听代码:

import telebot

# 替换成你的Bot Token(从BotFather那里获取)
bot = telebot.TeleBot("YOUR_BOT_TOKEN")

# 注册语音消息处理器
@bot.message_handler(content_types=['voice'])
def handle_voice_message(message):
    # 后续的语音处理逻辑都放在这里
    pass

# 启动Bot
bot.polling()

这段代码会让Bot开始监听所有用户发送的语音消息,接下来我们要获取语音文件本身。

2. 获取Telegram的语音文件

用户发送的语音消息里包含voice.file_id,我们可以用这个ID调用Telegram的getFile接口拿到文件的下载路径,然后把文件下载下来(或者直接传给Google API)。修改上面的处理器:

@bot.message_handler(content_types=['voice'])
def handle_voice_message(message):
    # 通过file_id获取文件信息
    file_info = bot.get_file(message.voice.file_id)
    # 下载语音文件到本地(存为voice.ogg,Telegram语音默认是OGG格式)
    downloaded_file = bot.download_file(file_info.file_path)
    with open("voice.ogg", 'wb') as f:
        f.write(downloaded_file)

不用担心格式问题,Google Speech-to-Text原生支持OGG_OPUS格式,直接用就行。

3. 用Google Speech-to-Text转语音为文本

接下来要调用Google的语音识别API。首先得准备好:

  • 安装Google客户端库:pip install google-cloud-speech
  • 配置Google服务账号密钥(把密钥文件路径设为环境变量GOOGLE_APPLICATION_CREDENTIALS)

然后写一个转文本的函数:

from google.cloud import speech_v1p1beta1 as speech

def transcribe_voice_to_text(file_path):
    client = speech.SpeechClient()

    # 读取本地语音文件
    with open(file_path, "rb") as audio_file:
        audio_content = audio_file.read()

    # 配置识别参数
    audio = speech.RecognitionAudio(content=audio_content)
    config = speech.RecognitionConfig(
        encoding=speech.RecognitionConfig.AudioEncoding.OGG_OPUS,
        sample_rate_hertz=48000,  # Telegram语音通常是48kHz采样率
        language_code="zh-CN",  # 替换成你需要识别的源语言,比如英文是en-US
    )

    # 调用API识别
    response = client.recognize(config=config, audio=audio)
    
    # 提取最准确的识别结果
    transcript = ""
    for result in response.results:
        transcript += result.alternatives[0].transcript
    return transcript

把这个函数集成到处理器里,就能拿到语音转成的文本了。

4. 文本翻译(按需实现)

如果需要把识别后的文本翻译成其他语言,比如中文转英文,可以用Google Translate API。同样先准备:

  • 安装客户端库:pip install google-cloud-translate
  • 确保服务账号有Translate API的权限

写翻译函数:

from google.cloud import translate_v2 as translate

def translate_text_content(text, target_lang="en"):
    translate_client = translate.Client()
    # 调用翻译API
    translation_result = translate_client.translate(text, target_language=target_lang)
    return translation_result['translatedText']

5. 把结果返回给用户

最后把所有步骤串起来,把识别和翻译的结果发回给用户:

@bot.message_handler(content_types=['voice'])
def handle_voice_message(message):
    try:
        # 下载语音文件
        file_info = bot.get_file(message.voice.file_id)
        downloaded_file = bot.download_file(file_info.file_path)
        with open("voice.ogg", 'wb') as f:
            f.write(downloaded_file)
        
        # 语音转文本
        original_text = transcribe_voice_to_text("voice.ogg")
        if not original_text:
            bot.reply_to(message, "抱歉,我没听清你说的内容😅")
            return
        
        # 翻译文本(如果不需要翻译可以删掉这步)
        translated_text = translate_text_content(original_text, target_lang="en")
        
        # 回复用户
        reply_msg = f"你说的是:{original_text}\n翻译结果:{translated_text}"
        bot.reply_to(message, reply_msg)
    except Exception as e:
        bot.reply_to(message, f"出错了:{str(e)}")

一些实用提示

  • 如果不想把文件存到本地,可以直接把downloaded_file的二进制内容传给Google API,不用写文件到磁盘。
  • Google API有免费额度,开发测试完全够用,正式使用需要留意计费规则。
  • 可以优化语言检测逻辑,比如让用户发送语音时附带目标语言,或者自动检测源语言。
  • 确保你的Bot在BotFather那里没有被限制接收语音消息(默认是允许的)。

内容的提问来源于stack exchange,提问作者Arman Fatahi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 03:42:28