You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

能否通过Python将预录音频文件传入Google Assistant SDK并构建API?

Answers to Your Google Assistant SDK Questions

Let's break down your questions one by one, with practical Python-focused solutions tailored to your use case.

1. Can I directly pass an audio file as input to the Google Assistant SDK?

Not exactly "directly" in the sense of passing a file path, but you can feed pre-recorded audio data to the SDK—so long as the audio meets strict format requirements:

  • 16kHz sampling rate
  • Mono (single channel)
  • Uncompressed PCM (typically saved as a .raw file)

The SDK expects audio input as a byte stream or raw audio bytes, not a file reference. So you’ll need to read your pre-recorded file into memory and pass those bytes to the SDK’s input methods.

2. Can I trigger a fixed pre-recorded command via a button/API endpoint?

Absolutely! This is totally feasible with the Google Assistant SDK, and Python is perfect for building the API endpoint you need. Here’s a step-by-step implementation plan plus a working code example:

Prerequisites

First, get your environment set up:

  • Enable the Google Assistant API in your Google Cloud Project, then create and download service account credentials (save as credentials.json).
  • Install required Python packages:
    pip install google-assistant-sdk[samples] flask google-auth
    
  • Convert your pre-recorded command audio to the correct format (if it’s not already). Use ffmpeg for this quick conversion:
    ffmpeg -i your-command.mp3 -acodec pcm_s16le -ac 1 -ar 16000 command.raw
    

Python API Endpoint Example (Using Flask)

This code creates a simple POST endpoint that triggers the Assistant with your pre-recorded audio when called—exactly what you need for a button-triggered workflow:

from flask import Flask, jsonify
import google.auth
from google.assistant.library import Assistant
from google.assistant.library.event import EventType
import os

# Initialize Flask app
app = Flask(__name__)

# Load pre-recorded audio once at startup (avoids reloading on every request)
PRE_RECORDED_AUDIO_PATH = "command.raw"
if not os.path.exists(PRE_RECORDED_AUDIO_PATH):
    raise FileNotFoundError(f"Audio file {PRE_RECORDED_AUDIO_PATH} not found. Check your conversion step!")

with open(PRE_RECORDED_AUDIO_PATH, "rb") as f:
    PRE_RECORDED_AUDIO = f.read()

def fetch_assistant_response():
    # Load Google credentials (set the GOOGLE_APPLICATION_CREDENTIALS env var to your credentials.json path)
    credentials, project_id = google.auth.default(scopes=[
        "https://www.googleapis.com/auth/assistant-sdk-prototype"
    ])
    
    with Assistant(credentials, project_id) as assistant:
        assistant.start_conversation()
        # Send the pre-recorded audio bytes to the Assistant
        assistant.send_audio(PRE_RECORDED_AUDIO)
        
        # Listen for and return the Assistant's response
        for event in assistant.start():
            if event.type == EventType.ON_RESPONSE_RECEIVED:
                response_text = event.args.get('text', 'No text response received')
                assistant.stop_conversation()
                return response_text
            elif event.type == EventType.ON_CONVERSATION_TERMINATED:
                assistant.stop_conversation()
                return "Conversation ended without a response"

@app.route('/trigger-command', methods=['POST'])
def trigger_command():
    try:
        response = fetch_assistant_response()
        return jsonify({
            "status": "success",
            "assistant_response": response
        })
    except Exception as e:
        return jsonify({
            "status": "error",
            "message": str(e)
        }), 500

if __name__ == '__main__':
    # Run the API endpoint (accessible at http://localhost:5000/trigger-command)
    app.run(host='0.0.0.0', port=5000, debug=True)

Key Tips for Production

  • Audio Format Check: Double-check your .raw file matches the 16kHz mono PCM requirement—incorrect format will make the Assistant ignore your input.
  • Authentication: Ensure the GOOGLE_APPLICATION_CREDENTIALS environment variable points to your credentials.json file before running the app.
  • Performance Optimization: For high-traffic use, reuse the Assistant instance instead of creating a new one per request (use a singleton pattern to reduce latency).
  • Error Handling: Expand the basic error handling to cover cases like missing credentials, audio read failures, or SDK timeouts.

内容的提问来源于stack exchange,提问作者joke4me

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:48:24