如何通过API自动获取Telegram Premium语音转文本结果?
可行方案说明
由于你已订阅Telegram Premium,完全可以通过Telethon调用官方API批量触发语音转写并直接获取Telegram原生转写结果,以下是具体实现方案:
核心前提
- 确保使用的Telegram账号已激活Premium(官方语音转写为Premium专属功能)
- 提前在Telegram开发者平台申请API ID和API Hash(创建应用即可获取,无需额外审核)
实现步骤与代码示例
1. 环境准备
安装Telethon依赖:
pip install telethon
2. 批量转写核心代码
from telethon import TelegramClient from telethon.tl.types import Voice from telethon.tl.functions.messages import RequestTranscriptionRequest import asyncio import json # 替换为你的实际参数 API_ID = 123456 # 你的API ID API_HASH = "abcdef1234567890abcdef1234567890" # 你的API Hash SESSION_NAME = "premium_transcribe_session" # 会话文件名,首次运行会生成 async def fetch_transcription(client, chat_id, msg_id): """单条语音消息的转写触发与结果获取""" msg = await client.get_messages(chat_id, ids=msg_id) if not isinstance(msg.media, Voice): return {"message_id": msg_id, "error": "Not a voice message"} # 触发官方转写请求 try: await client(RequestTranscriptionRequest( peer=chat_id, msg_id=msg_id, lang_code=None # 留空自动检测语言,可指定如"zh"、"en" )) # 轮询等待转写完成(Telegram转写通常需1-5秒) for _ in range(10): # 最多等待10秒 updated_msg = await client.get_messages(chat_id, ids=msg_id) if hasattr(updated_msg.media, "transcription") and updated_msg.media.transcription: return { "message_id": msg_id, "transcription": updated_msg.media.transcription } await asyncio.sleep(1) return {"message_id": msg_id, "error": "Transcription timeout"} except Exception as e: return {"message_id": msg_id, "error": str(e)} async def batch_process(chat_id, message_ids): """批量处理语音消息转写""" async with TelegramClient(SESSION_NAME, API_ID, API_HASH) as client: await client.start() # 首次运行会要求扫码登录 results = [] for idx, msg_id in enumerate(message_ids): print(f"Processing message {idx+1}/{len(message_ids)}") res = await fetch_transcription(client, chat_id, msg_id) results.append(res) # 加延迟避免触发API速率限制 await asyncio.sleep(1.5) # 保存结果到本地文件 with open("telegram_transcriptions.json", "w", encoding="utf-8") as f: json.dump(results, f, ensure_ascii=False, indent=2) print("Batch processing completed, results saved to telegram_transcriptions.json") if __name__ == "__main__": # 替换为你的聊天ID(可通过Telethon的get_entity方法获取)和目标消息ID列表 TARGET_CHAT_ID = -1001234567890 # 超级群/频道的ID,前缀为-100 TARGET_MESSAGE_IDS = list(range(1, 5001)) # 示例:处理ID1到5000的消息 asyncio.run(batch_process(TARGET_CHAT_ID, TARGET_MESSAGE_IDS))
关键注意事项
- 速率限制:Telegram对API请求有严格的频率限制,批量处理时必须添加1-2秒的延迟,否则可能触发临时封禁
- 转写失败处理:部分低质量、非人类语音或过长的音频可能无法完成转写,代码中已包含超时和异常捕获逻辑
- 会话安全:会话文件(
premium_transcribe_session.session)包含账号登录信息,请勿泄露给他人 - 语言指定:若需强制转写特定语言,可将
lang_code参数设置为对应语言代码(如"zh-CN"、"en-US")
内容的提问来源于stack exchange,提问作者OnTheSky9
相关产品推荐
相关产品推荐

