You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在不生成MP3文件的情况下获取pyttsx3生成的语音时长?

不生成语音文件获取TTS时长的高精度方法

针对你使用pyttsx3、语速175、指定男声的场景,要达到10~100毫秒的精度,推荐以下两种可行方案:

方案1:利用空设备合成+事件回调(最可靠,精度最高)

pyttsx3支持将合成内容输出到系统空设备(不会生成实际文件),同时通过事件回调监听合成的开始与结束时间,直接计算实际时长。这种方法基于真实合成过程,精度完全满足要求。

示例代码:

import pyttsx3
import time

total_duration = 0.0
utterance_start = None

def track_start(name):
    global utterance_start
    utterance_start = time.perf_counter()

def track_end(name, completed):
    global total_duration
    if completed and utterance_start:
        total_duration += time.perf_counter() - utterance_start

# 初始化引擎并配置参数
engine = pyttsx3.init()
engine.setProperty('rate', 175)
target_voice = engine.getProperty('voices')[0]
engine.setProperty('voice', target_voice.id)

# 注册时长跟踪回调
engine.connect('started-utterance', track_start)
engine.connect('finished-utterance', track_end)

# 要计算时长的目标文本
input_text = "Hello, this is a test sentence for duration calculation."

# 输出到空设备(Windows用'nul',Mac/Linux用'/dev/null')
engine.save_to_file(input_text, 'nul')
engine.runAndWait()

# 输出结果(转成毫秒,保留两位小数)
print(f"文本语音时长:{total_duration * 1000:.2f} 毫秒")

方案2:音素级预校准计算(离线无合成,精度依赖校准)

如果完全不想触发合成流程,可以先针对目标男声做一次音素时长校准,之后通过文本转音素+时长累加的方式估算:

  1. 生成包含所有常见音素的校准文本,用方案1的方法获取每个音素的实际时长,建立音素-时长映射表。
  2. 将目标文本转成音素序列(可借助CMU词典等工具),累加对应音素的时长,再根据语速175做比例调整。

示例代码(假设已完成校准):

from nltk.corpus import cmudict

# 预校准的音素时长字典(示例值,需实际校准)
phoneme_durations = {
    'AA': 0.120, 'AE': 0.105, 'AH': 0.098,
    # 补充所有需要的音素时长...
}

# 文本转音素(基于CMU词典)
cmu_dict = cmudict.dict()
def text_to_phonemes(text):
    phonemes = []
    for word in text.lower().split():
        if word in cmu_dict:
            phonemes.extend(cmu_dict[word][0])
    return phonemes

# 计算时长
def get_tts_duration(text, rate=175):
    # 基准语速设为pyttsx3默认的200,根据目标语速调整比例
    rate_adjust = 200 / rate
    phonemes = text_to_phonemes(text)
    total_sec = sum(phoneme_durations[p] for p in phonemes) * rate_adjust
    return total_sec * 1000

# 测试
input_text = "Hello, this is a test sentence."
print(f"估算时长:{get_tts_duration(input_text):.2f} 毫秒")

注意事项

  • 方案1的空设备合成不会产生实际文件,只是利用引擎的合成流程获取真实时长,精度能稳定在10毫秒以内。
  • 方案2的精度取决于校准的准确性,建议用包含各种音素、音节结构的长文本做校准,减少误差。

内容的提问来源于stack exchange,提问作者ThatOneGuy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 21:10:30