能否仅使用Amazon Lex实现语音转文本并将文本传入Lambda函数?
Absolutely! This workflow is fully supported out of the box with Amazon Lex, and it aligns perfectly with your goal of capturing full user speech, converting it to text, and processing that text in a custom Lambda function. Here's a step-by-step breakdown to set it up:
1. Enable Speech Input in Your Lex Bot
Amazon Lex leverages Amazon Transcribe's automatic speech recognition (ASR) under the hood to convert voice input to text—you don’t need to manually integrate Transcribe for this core functionality. When configuring your bot:
- In the Lex Console, ensure Voice is selected as an interaction mode (this is enabled by default for most bot templates).
- You can tweak speech recognition settings (like language support, confidence thresholds) in the bot’s Settings > Speech tab if needed. Lex will automatically capture the full user speech, transcribe it to text, and include that text in the request it sends to Lambda.
2. Link Your Lambda Function as a Code Hook
To route the transcribed text to your Lambda function:
- Create an Intent in your Lex Bot (e.g.,
ProcessUserVoiceInput—this acts as the trigger for your workflow). - Under the intent’s Fulfillment section, select your Lambda function as the code hook. Lex will send a request to this Lambda every time the intent is triggered by a user’s voice input.
- Make sure your Lambda has the correct permissions: Lex needs permission to invoke your function, which you can set up via IAM (Lex will even prompt you to create the necessary role when linking the function initially).
3. Access and Process the Transcribed Text in Lambda
When Lex sends a request to your Lambda, the full transcribed speech text lives in the inputTranscript field of the event object. Here’s a quick Python example to demonstrate accessing and processing it:
def lambda_handler(event, context): # Pull the full transcribed text from Lex's request user_transcript = event["inputTranscript"] # Your custom processing logic goes here (e.g., keyword extraction, sentiment analysis) processed_output = f"Successfully processed your input: {user_transcript}" # Return a valid response to Lex to close or continue the conversation return { "sessionState": { "dialogAction": {"type": "Close"}, "intent": {"name": "ProcessUserVoiceInput", "state": "Fulfilled"} }, "messages": [{"contentType": "PlainText", "content": processed_output}] }
Quick Notes to Keep in Mind
- If you ever need access to the raw audio file (not just the transcribed text), Lex doesn’t send this by default—you’d need to integrate with Amazon Connect or build a custom frontend to capture audio first. But for your stated goal, Lex’s built-in ASR is more than enough.
- Lex supports multiple languages for speech recognition—double-check your bot’s language setting matches the user’s input language for accurate transcription.
- You can adjust the confidence threshold in Lex’s speech settings to filter out low-confidence transcriptions before they’re sent to Lambda.
内容的提问来源于stack exchange,提问作者Mohit Sarpal

