能否通过Python将预录音频文件传入Google Assistant SDK并构建API?
Let's break down your questions one by one, with practical Python-focused solutions tailored to your use case.
1. Can I directly pass an audio file as input to the Google Assistant SDK?
Not exactly "directly" in the sense of passing a file path, but you can feed pre-recorded audio data to the SDK—so long as the audio meets strict format requirements:
- 16kHz sampling rate
- Mono (single channel)
- Uncompressed PCM (typically saved as a
.rawfile)
The SDK expects audio input as a byte stream or raw audio bytes, not a file reference. So you’ll need to read your pre-recorded file into memory and pass those bytes to the SDK’s input methods.
2. Can I trigger a fixed pre-recorded command via a button/API endpoint?
Absolutely! This is totally feasible with the Google Assistant SDK, and Python is perfect for building the API endpoint you need. Here’s a step-by-step implementation plan plus a working code example:
Prerequisites
First, get your environment set up:
- Enable the Google Assistant API in your Google Cloud Project, then create and download service account credentials (save as
credentials.json). - Install required Python packages:
pip install google-assistant-sdk[samples] flask google-auth - Convert your pre-recorded command audio to the correct format (if it’s not already). Use
ffmpegfor this quick conversion:ffmpeg -i your-command.mp3 -acodec pcm_s16le -ac 1 -ar 16000 command.raw
Python API Endpoint Example (Using Flask)
This code creates a simple POST endpoint that triggers the Assistant with your pre-recorded audio when called—exactly what you need for a button-triggered workflow:
from flask import Flask, jsonify import google.auth from google.assistant.library import Assistant from google.assistant.library.event import EventType import os # Initialize Flask app app = Flask(__name__) # Load pre-recorded audio once at startup (avoids reloading on every request) PRE_RECORDED_AUDIO_PATH = "command.raw" if not os.path.exists(PRE_RECORDED_AUDIO_PATH): raise FileNotFoundError(f"Audio file {PRE_RECORDED_AUDIO_PATH} not found. Check your conversion step!") with open(PRE_RECORDED_AUDIO_PATH, "rb") as f: PRE_RECORDED_AUDIO = f.read() def fetch_assistant_response(): # Load Google credentials (set the GOOGLE_APPLICATION_CREDENTIALS env var to your credentials.json path) credentials, project_id = google.auth.default(scopes=[ "https://www.googleapis.com/auth/assistant-sdk-prototype" ]) with Assistant(credentials, project_id) as assistant: assistant.start_conversation() # Send the pre-recorded audio bytes to the Assistant assistant.send_audio(PRE_RECORDED_AUDIO) # Listen for and return the Assistant's response for event in assistant.start(): if event.type == EventType.ON_RESPONSE_RECEIVED: response_text = event.args.get('text', 'No text response received') assistant.stop_conversation() return response_text elif event.type == EventType.ON_CONVERSATION_TERMINATED: assistant.stop_conversation() return "Conversation ended without a response" @app.route('/trigger-command', methods=['POST']) def trigger_command(): try: response = fetch_assistant_response() return jsonify({ "status": "success", "assistant_response": response }) except Exception as e: return jsonify({ "status": "error", "message": str(e) }), 500 if __name__ == '__main__': # Run the API endpoint (accessible at http://localhost:5000/trigger-command) app.run(host='0.0.0.0', port=5000, debug=True)
Key Tips for Production
- Audio Format Check: Double-check your
.rawfile matches the 16kHz mono PCM requirement—incorrect format will make the Assistant ignore your input. - Authentication: Ensure the
GOOGLE_APPLICATION_CREDENTIALSenvironment variable points to yourcredentials.jsonfile before running the app. - Performance Optimization: For high-traffic use, reuse the Assistant instance instead of creating a new one per request (use a singleton pattern to reduce latency).
- Error Handling: Expand the basic error handling to cover cases like missing credentials, audio read failures, or SDK timeouts.
内容的提问来源于stack exchange,提问作者joke4me

