如何用Picovoice Leopard与PyAudio转录原始音频数据?
解决Picovoice Leopard调用时的TypeError问题
问题原因
你的get_audio_data()返回的是bytes对象的列表,但Leopard的process方法需要的是16位整数数组(int16),而非原始bytes块的集合,这导致了类型错误。
修正方案
需要将PyAudio读取的bytes数据转换为Leopard要求的int16数组,步骤如下:
- 首先安装numpy(用于音频数据转换):
pip install numpy
- 修改
get_audio_data函数,合并bytes并转换为int16数组,同时正确释放音频资源:
import numpy as np import pyaudio AUDIO_FORMAT = pyaudio.paInt16 AUDIO_CHANNELS = 1 SAMPLE_RATE = 16000 CHUNK_SIZE = 512 def get_audio_data(): audio = pyaudio.PyAudio() stream = audio.open(format=AUDIO_FORMAT, channels=AUDIO_CHANNELS, rate=SAMPLE_RATE, input=True, frames_per_buffer=CHUNK_SIZE) frames = [] while True: data = stream.read(CHUNK_SIZE) frames.append(data) # 替换为你的停止条件,例如录制5秒: # if len(frames) * CHUNK_SIZE / SAMPLE_RATE >= 5: # break # 合并所有bytes块为一个完整的bytes对象 audio_bytes = b''.join(frames) # 将bytes转换为int16数组(paInt16对应numpy.int16类型) audio_array = np.frombuffer(audio_bytes, dtype=np.int16) # 释放PyAudio资源 stream.stop_stream() stream.close() audio.terminate() return audio_array
- 正常调用Leopard的process方法:
transcript, words = leopard.process(get_audio_data()) print(transcript) for word in words: print( "{word=\"%s\" start_sec=%.2f end_sec=%.2f confidence=%.2f}" % (word.word, word.start_sec, word.end_sec, word.confidence))
额外注意事项
- 确保Leopard初始化时的采样率与录音采样率一致(你设置的16000是Leopard的默认值,无需额外调整,若修改采样率需同步两边参数)。
- 不要忘记填写你的Picovoice Access Key初始化Leopard:
from pvleopard import Leopard leopard = Leopard(access_key="你的AccessKey", sample_rate=SAMPLE_RATE)
内容的提问来源于stack exchange,提问作者Almogbb
相关产品推荐
相关产品推荐

