You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java版Google Speech-to-Text流式音频转写结果不稳定问题求助

我来帮你排查下这个流式语音识别不稳定的问题,看起来你的代码里有几个关键细节没处理好,导致结果时好时坏:

核心问题分析

  1. onComplete()调用时机错误
    你在try块外面手动调用了responseObserver.onComplete(),但实际上当SpeechClient通过try-with-resources自动关闭时,gRPC框架已经会触发onComplete()回调了。手动重复调用会导致两种情况:要么回调还没处理完所有API响应就提前执行,要么覆盖了之前的有效结果,最终返回空字符串。

  2. 音频流未正确关闭
    当你检测到用户停止说话后,直接关闭了麦克风输入,但没有告诉Google API“音频已经发送完毕”。服务器可能还在等待更多数据,不会返回最终的识别结果,自然就没有转录文本了。

  3. 线程同步问题
    识别响应的处理是在回调线程里完成的,而主线程在try块结束后直接读取SPEECH_TO_TEXT_ANSWER,很可能出现主线程先执行完,回调线程还没处理完API响应的情况,这时候读取到的就是初始的空字符串。

修复后的代码示例

针对上面的问题,我调整了你的代码,主要加入了线程同步机制、正确的流关闭逻辑,去掉了多余的onComplete()调用:

import java.util.concurrent.CountDownLatch;
import java.util.concurrent.atomic.AtomicReference;

public static String streamingMicRecognize(String language) throws Exception {
    final CountDownLatch completionLatch = new CountDownLatch(1);
    final AtomicReference<String> finalTranscript = new AtomicReference<>("");
    ResponseObserver<StreamingRecognizeResponse> responseObserver = null;

    try (SpeechClient client = SpeechClient.create()) {
        responseObserver = new ResponseObserver<StreamingRecognizeResponse>() {
            ArrayList<StreamingRecognizeResponse> responses = new ArrayList<>();

            @Override
            public void onStart(StreamController controller) {}

            @Override
            public void onResponse(StreamingRecognizeResponse response) {
                responses.add(response);
            }

            @Override
            public void onComplete() {
                StringBuilder transcriptBuilder = new StringBuilder();
                for (StreamingRecognizeResponse response : responses) {
                    if (!response.getResultsList().isEmpty()) {
                        StreamingRecognitionResult result = response.getResultsList().get(0);
                        if (!result.getAlternativesList().isEmpty()) {
                            SpeechRecognitionAlternative alternative = result.getAlternativesList().get(0);
                            System.out.printf("Transcript : %s\n", alternative.getTranscript());
                            transcriptBuilder.append(alternative.getTranscript());
                        }
                    }
                }
                finalTranscript.set(transcriptBuilder.toString());
                completionLatch.countDown(); // 通知主线程识别完成
            }

            @Override
            public void onError(Throwable t) {
                System.out.println("识别过程出错:" + t.getMessage());
                completionLatch.countDown(); // 出错时也通知主线程,避免无限等待
            }
        };

        ClientStream<StreamingRecognizeRequest> clientStream = client.streamingRecognizeCallable().splitCall(responseObserver);

        // 初始化识别配置(这部分和你原来的代码一致)
        RecognitionConfig recognitionConfig = RecognitionConfig.newBuilder()
                .setEncoding(RecognitionConfig.AudioEncoding.LINEAR16)
                .setLanguageCode(language)
                .setSampleRateHertz(16000)
                .build();
        StreamingRecognitionConfig streamingRecognitionConfig = StreamingRecognitionConfig.newBuilder().setConfig(recognitionConfig).build();
        StreamingRecognizeRequest request = StreamingRecognizeRequest.newBuilder()
                .setStreamingConfig(streamingRecognitionConfig)
                .build();
        clientStream.send(request);

        // 麦克风音频读取逻辑(这部分基础不变,仅调整音频发送细节)
        AudioFormat audioFormat = new AudioFormat(16000, 16, 1, true, false);
        DataLine.Info targetInfo = new DataLine.Info(TargetDataLine.class, audioFormat);
        if (!AudioSystem.isLineSupported(targetInfo)) {
            System.out.println("Microphone not supported");
            System.exit(0);
        }
        TargetDataLine targetDataLine = (TargetDataLine) AudioSystem.getLine(targetInfo);
        targetDataLine.open(audioFormat);
        targetDataLine.start();
        System.out.println("Start speaking");
        playMP3("beep-07.mp3");
        long startTime = System.currentTimeMillis();
        AudioInputStream audio = new AudioInputStream(targetDataLine);

        long estimatedTime = 0, estimatedTimeStoppedSpeaking = 0, startStopSpeaking = 0;
        int currentSoundLevel = 0;
        Boolean hasSpoken = false;

        while (true) {
            estimatedTime = System.currentTimeMillis() - startTime;
            byte[] data = new byte[6400];
            int bytesRead = audio.read(data); // 记录实际读取的字节数
            if (bytesRead <= 0) break;

            currentSoundLevel = calculateRMSLevel(data);
            System.out.println(currentSoundLevel);

            if (currentSoundLevel > 20) {
                estimatedTimeStoppedSpeaking = 0;
                startStopSpeaking = 0;
                hasSpoken = true;
            } else {
                if (startStopSpeaking == 0) {
                    startStopSpeaking = System.currentTimeMillis();
                }
                estimatedTimeStoppedSpeaking = System.currentTimeMillis() - startStopSpeaking;
            }

            if ((estimatedTime > 15000) || (estimatedTimeStoppedSpeaking > 1000 && hasSpoken)) {
                playMP3("beep-07.mp3");
                System.out.println("Stop speaking.");
                targetDataLine.stop();
                targetDataLine.drain();
                targetDataLine.close();
                clientStream.closeSend(); // 关键:告诉服务器音频发送完毕
                break;
            }

            // 只发送实际读取到的音频数据,避免空数据
            request = StreamingRecognizeRequest.newBuilder()
                    .setAudioContent(ByteString.copyFrom(data, 0, bytesRead))
                    .build();
            clientStream.send(request);
        }

        completionLatch.await(); // 等待回调线程处理完所有响应
    } catch (Exception e) {
        System.out.println("程序异常:" + e.getMessage());
    }

    return finalTranscript.get();
}

额外优化建议

  • 调整声音阈值:你设置的currentSoundLevel > 20可能需要根据实际环境调整,比如在安静环境下可以降低阈值,避免漏判用户说话;在嘈杂环境下提高阈值,避免误判背景音为说话声。
  • 测试音频读取:可以把读取到的音频数据保存成临时文件,验证是否和你单独录制的音频一致,确保麦克风输入的音频没有损坏。

内容的提问来源于stack exchange,提问作者g123

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 07:09:23