You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将AWS Transcribe Stream与React应用集成实现实时语音转写

AWS Transcribe Streaming 实时语音转写 React 集成方案

你的代码核心问题在于没有实现真正的实时音频流传输——当前是等录制停止后才把所有音频一次性发给Transcribe,而且MediaRecorder默认输出的格式也不符合AWS要求的PCM规范。以下是修正后的完整实现:

关键问题分析

  • MediaRecorder无法直接输出AWS要求的原始PCM格式,且无法实时分块发送音频
  • 音频流仅在录制结束后发送一次,不符合Transcribe Streaming的实时交互要求
  • 缺少对转录中间结果(Partial Transcript)的处理,无法实时更新文本

修正后的代码实现

首先安装依赖:

npm install microphone-stream @aws-sdk/client-transcribe-streaming

完整组件代码:

import React, { useState, useRef, useEffect } from 'react';
import { TranscribeStreamingClient, StartStreamTranscriptionCommand } from '@aws-sdk/client-transcribe-streaming';
import MicrophoneStream from 'microphone-stream';

interface TranscriptResult {
    Alternatives: { Items: { Content: string }[]; Transcript: string }[];
    IsPartial: boolean;
}

const TranscriptionComponent: React.FC = () => {
    const [transcript, setTranscript] = useState<string>('');
    const [isRecording, setIsRecording] = useState<boolean>(false);
    const [error, setError] = useState<string | null>(null);

    const micStreamRef = useRef<MicrophoneStream | null>(null);
    const clientRef = useRef<TranscribeStreamingClient | null>(null);

    useEffect(() => {
        // 初始化AWS Transcribe客户端
        clientRef.current = new TranscribeStreamingClient({
            region: '你的AWS区域', // 例如us-east-1
            credentials: {
                accessKeyId: '你的IAM Access Key',
                secretAccessKey: '你的IAM Secret Key',
            },
        });

        return () => {
            stopRecording();
        };
    }, []);

    const startRecording = async () => {
        try {
            setIsRecording(true);
            setError(null);
            setTranscript('');

            // 初始化麦克风流,获取PCM格式音频
            const micStream = new MicrophoneStream();
            micStreamRef.current = micStream;

            // 请求麦克风权限
            const stream = await navigator.mediaDevices.getUserMedia({ audio: true });
            micStream.setStream(stream);

            // 创建音频流生成器,按AWS要求的格式分块发送
            const audioStream = async function* () {
                for await (const chunk of micStream) {
                    // 将Float32Array转换为AWS要求的Uint8Array(PCM 16位)
                    const pcm16 = MicrophoneStream.toRaw(chunk);
                    yield { AudioEvent: { AudioChunk: pcm16 } };
                }
            };

            // 发送转录请求
            const command = new StartStreamTranscriptionCommand({
                LanguageCode: 'zh-CN', // 根据需求修改,例如en-US
                MediaEncoding: 'pcm',
                MediaSampleRateHertz: 44100,
                AudioStream: audioStream(),
            });

            const client = clientRef.current;
            if (!client) return;

            const response = await client.send(command);
            handleTranscriptionResponse(response);

        } catch (err) {
            console.error('麦克风或转录服务错误:', err);
            setError('无法启动转录,请检查麦克风权限和AWS配置');
            setIsRecording(false);
        }
    };

    const handleTranscriptionResponse = async (response: any) => {
        try {
            for await (const event of response.TranscriptResultStream) {
                if (event.TranscriptEvent) {
                    const results: TranscriptResult[] = event.TranscriptEvent.Transcript.Results;
                    results.forEach(result => {
                        // 处理部分结果(实时更新)和最终结果
                        if (result.Alternatives.length > 0) {
                            const currentTranscript = result.Alternatives[0].Transcript;
                            if (result.IsPartial) {
                                // 临时替换当前句的部分结果
                                setTranscript(prev => {
                                    const lastSpaceIndex = prev.lastIndexOf(' ');
                                    return lastSpaceIndex !== -1 ? prev.slice(0, lastSpaceIndex + 1) + currentTranscript : currentTranscript;
                                });
                            } else {
                                // 追加最终结果
                                setTranscript(prev => `${prev} ${currentTranscript}`.trim());
                            }
                        }
                    });
                }
            }
        } catch (err) {
            console.error('转录结果处理错误:', err);
            setError('转录过程出错,请重试');
        } finally {
            stopRecording();
        }
    };

    const stopRecording = () => {
        if (micStreamRef.current) {
            micStreamRef.current.stop();
            micStreamRef.current = null;
        }
        setIsRecording(false);
    };

    return (
        <div className="container">
            <h1>实时语音转写</h1>
            <hr />
            {error && <div id="error" className="isa_error">{error}</div>}
            <textarea
                id="transcript"
                placeholder="点击开始,然后说话"
                rows={8}
                readOnly
                value={transcript}
                style={{ width: '100%', padding: '10px', fontSize: '16px' }}
            />
            <div className="row" style={{ marginTop: '20px' }}>
                <button
                    id="start-button"
                    className="button-xl"
                    onClick={startRecording}
                    disabled={isRecording}
                    style={{ padding: '10px 20px', marginRight: '10px', fontSize: '16px' }}
                >
                    🎤 开始转写
                </button>
                <button
                    id="stop-button"
                    className="button-xl"
                    onClick={stopRecording}
                    disabled={!isRecording}
                    style={{ padding: '10px 20px', fontSize: '16px' }}
                >
                    ⏹️ 停止转写
                </button>
            </div>
        </div>
    );
};

export default TranscriptionComponent;

核心修改说明

  1. 替换音频捕获方式:使用microphone-stream直接获取PCM格式的音频流,无需等待录制结束,支持实时分块发送
  2. 实时流传输:通过异步生成器持续向AWS Transcribe发送音频块,符合Streaming API的要求
  3. 处理部分转录结果:识别IsPartial字段,实时更新当前正在说话的内容,提升用户体验
  4. 格式转换:将麦克风输出的Float32Array转换为AWS要求的16位PCM Uint8Array

注意事项

  • 确保IAM用户拥有transcribe:StartStreamTranscription权限
  • 音频采样率需与麦克风实际采样率匹配(通常为44100或16000)
  • 生产环境中不要硬编码AWS凭证,建议使用Cognito或其他安全方式获取临时凭证

内容的提问来源于stack exchange,提问作者Sandun Tharaka

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 13:23:12