You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Picovoice Leopard与PyAudio转录原始音频数据?

解决Picovoice Leopard调用时的TypeError问题

问题原因

你的get_audio_data()返回的是bytes对象的列表,但Leopard的process方法需要的是16位整数数组(int16),而非原始bytes块的集合,这导致了类型错误。

修正方案

需要将PyAudio读取的bytes数据转换为Leopard要求的int16数组,步骤如下:

  1. 首先安装numpy(用于音频数据转换):
pip install numpy
  1. 修改get_audio_data函数,合并bytes并转换为int16数组,同时正确释放音频资源:
import numpy as np
import pyaudio

AUDIO_FORMAT = pyaudio.paInt16 
AUDIO_CHANNELS = 1  
SAMPLE_RATE = 16000 
CHUNK_SIZE = 512

def get_audio_data():
    audio = pyaudio.PyAudio()
    stream = audio.open(format=AUDIO_FORMAT, channels=AUDIO_CHANNELS,
                            rate=SAMPLE_RATE, input=True,
                            frames_per_buffer=CHUNK_SIZE)
    frames = []
    while True:
        data = stream.read(CHUNK_SIZE)
        frames.append(data)
        # 替换为你的停止条件,例如录制5秒:
        # if len(frames) * CHUNK_SIZE / SAMPLE_RATE >= 5:
        #     break
    # 合并所有bytes块为一个完整的bytes对象
    audio_bytes = b''.join(frames)
    # 将bytes转换为int16数组(paInt16对应numpy.int16类型)
    audio_array = np.frombuffer(audio_bytes, dtype=np.int16)
    
    # 释放PyAudio资源
    stream.stop_stream()
    stream.close()
    audio.terminate()
    
    return audio_array
  1. 正常调用Leopard的process方法:
transcript, words = leopard.process(get_audio_data())
print(transcript)
for word in words:
    print(
      "{word=\"%s\" start_sec=%.2f end_sec=%.2f confidence=%.2f}"
      % (word.word, word.start_sec, word.end_sec, word.confidence))

额外注意事项

  • 确保Leopard初始化时的采样率与录音采样率一致(你设置的16000是Leopard的默认值,无需额外调整,若修改采样率需同步两边参数)。
  • 不要忘记填写你的Picovoice Access Key初始化Leopard:
from pvleopard import Leopard

leopard = Leopard(access_key="你的AccessKey", sample_rate=SAMPLE_RATE)

内容的提问来源于stack exchange,提问作者Almogbb

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 20:48:18