You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java调用Python音频识别脚本时卡在Vosk模型初始化阶段

问题:Java调用Python音频处理脚本时停滞在模型构建环节

单独运行Python音频处理脚本可正常执行,但通过Java构造函数调用时,程序输出停留在打印===> Build the model and recognizer objects. This will take a few minutes.这一行,后续标记为print("1")的代码及之后逻辑均未运行。已尝试用Python的os模块将模型相对路径改为绝对路径,但程序未抛出任何错误,仅停滞在此环节。


相关代码

ccPage.java

ccPage() throws IOException {
setLayout(null);

    getContentPane().setBackground(Color.WHITE);
    setSize(800, 480);
    setVisible(true);
    setLocation(350, 200);

    Path pythonScriptPath = Paths.get("").toAbsolutePath().resolve(Paths.get("src/software/pythonFiles/audioProcess.py"));

    // Create a new Python interpreter
    try(Interpreter py = new SubInterpreter()){
        // Execute the Python script
        py.runScript(pythonScriptPath.toString());

        // Get the variable `mystring` from the Python script
        String mystring = py.getValue("mystring").toString();

        // Create a JLabel to display the value of `mystring`
        mystringLabel = new JLabel(mystring);
        mystringLabel.setBounds(10, 10, 200, 20);
        add(mystringLabel);

        // Continuously fetch the updated value of `mystring` from the Python script
        while (true) {
            // Get the updated value of `mystring` from the Python script
            mystring = py.getValue("mystring").toString();

            // Update the JLabel to display the updated value of `mystring`
            mystringLabel.setText(mystring);

            // Sleep for 1 second
            try {
                Thread.sleep(1000);
            } catch (InterruptedException e) {
                e.printStackTrace();
            }
        }
    }
    catch (JepException e) {
        e.printStackTrace();
    }
}

audioProcess.py

import os
import queue

from pathlib import Path
import sounddevice as sd
from vosk import Model, KaldiRecognizer
import sys
import json

import subprocess                  #Test line!
pipe = subprocess.PIPE             #Test line!

'''This script processes audio input from the microphone and displays the transcribed text.'''

# list all audio devices known to your system
print("Display input/output devices")
print(sd.query_devices())

# get the samplerate - this is needed by the Kaldi recognizer
device_info = sd.query_devices(sd.default.device[0], 'input')
samplerate = int(device_info['default_samplerate'])

# display the default input device
print("===> Initial Default Device Number:{} Description: {}".format(sd.default.device[0], device_info))

# setup queue and callback function
q = queue.Queue()

myString = ""

def recordCallback(indata, frames, time, status):
    if status:
        print(status, file=sys.stderr)
    q.put(bytes(indata))


# build the model and recognizer objects.
print("===> Build the model and recognizer objects.  This will take a few minutes.")
model_file: Path = Path(os.path.abspath("vosk-model-small-en-us-0.15"))
print("1")
model = Model(model_file.as_posix())
recognizer = KaldiRecognizer(model, samplerate)

print("===> Begin recording. Press Ctrl+C to stop the recording ")
try:
    with sd.RawInputStream(dtype='int16',
                           channels=1,
                           callback=recordCallback,
                           blocksize=100,
                           samplerate=samplerate,
                           extra_settings=pipe
                           ):
        # made changes that update the variable to be printed...
        while True:
            data = q.get()
            if recognizer.AcceptWaveform(data):
                recognizerResult = recognizer.Result()
                # convert the recognizerResult string into a dictionary
                resultDict = json.loads(recognizerResult)
                if not resultDict.get("text", "") == "":
                    print(recognizerResult)
                    myString = recognizerResult
                else:
                    print("no input sound")
                    myString = ""

except KeyboardInterrupt:
    print('===> Finished Recording')
except Exception as e:
    print(str(e))

解决方案

1. 修复sounddevice初始化错误

Python脚本中sd.RawInputStream的extra_settings=pipe参数完全错误,extra_settings需要平台特定的音频配置对象(如Windows的ASIO设置),直接传入subprocess.PIPE会导致音频流初始化阻塞,这是卡住的核心原因。删除该行,修改后的代码:

with sd.RawInputStream(dtype='int16',
                       channels=1,
                       callback=recordCallback,
                       blocksize=100,
                       samplerate=samplerate
                       ):

2. 确认模型路径的有效性

在Python脚本中添加路径打印,验证Java环境下模型路径是否正确:

model_file: Path = Path(os.path.abspath("vosk-model-small-en-us-0.15"))
print(f"模型绝对路径:{model_file.as_posix()}")  # 新增该行
print("1")
model = Model(model_file.as_posix())

运行Java程序后查看输出路径,确认该路径存在且包含完整的Vosk模型文件(如am/final.mdl、graph/phones/word_boundary.int等)。若路径错误,直接替换为模型的绝对路径字符串(如Path("D:/project/vosk-model-small-en-us-0.15"))。

3. 调整Python脚本的执行逻辑

原Python脚本末尾的无限循环会导致Java的py.runScript调用永远无法返回,后续获取myString的代码无法执行。将音频处理逻辑放到后台线程,让脚本快速执行完成:

# 在脚本末尾添加线程逻辑
import threading

def audio_processing_loop():
    print("===> Begin recording. Press Ctrl+C to stop the recording ")
    try:
        with sd.RawInputStream(dtype='int16',
                               channels=1,
                               callback=recordCallback,
                               blocksize=100,
                               samplerate=samplerate
                               ):
            while True:
                data = q.get()
                if recognizer.AcceptWaveform(data):
                    recognizerResult = recognizer.Result()
                    resultDict = json.loads(recognizerResult)
                    if not resultDict.get("text", "") == "":
                        print(recognizerResult)
                        global myString  # 声明全局变量
                        myString = recognizerResult
                    else:
                        print("no input sound")
                        myString = ""
    except Exception as e:
        print(str(e))

# 启动后台守护线程执行音频处理
threading.Thread(target=audio_processing_loop, daemon=True).start()

这样Python脚本执行时会立即启动后台线程,Java的runScript调用能快速返回,正常获取到myString变量并进入更新循环。

4. 检查Java进程权限

确保Java进程有访问模型文件夹和音频设备的权限:

  • Windows下:右键以管理员身份运行Java程序,避免权限不足导致模型加载失败。
  • Linux/macOS下:确保Java进程所在用户有权限读取模型文件夹,且有权限访问麦克风(通过pactl等工具检查音频权限)。

内容的提问来源于stack exchange,提问作者Parth Verma

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 05:06:00