You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过CLI使用Google Assistant SDK+Dialogflow?自制AI音箱技术问询

可以在CLI中调用Dialogflow意图,实现类似pushtotalk.py的功能

当然没问题!你完全能在Ubuntu 16.04的命令行环境下直接调用Dialogflow的意图,不用局限于Actions on Google的网页模拟器。下面是适配你环境的具体实现方案:

1. 安装Dialogflow Python SDK

首先,安装Dialogflow的官方Python客户端库,这是CLI调用的基础:

pip install dialogflow

如果你的系统默认Python版本有冲突,可以使用虚拟环境,或者加上--user参数避免全局安装冲突:

pip install --user dialogflow

2. 获取Dialogflow服务账号密钥

要调用Dialogflow的API,必须通过服务账号完成认证:

  • 打开Dialogflow控制台,进入你的目标Agent
  • 点击左侧菜单栏的设置(齿轮图标)
  • 在服务账号区域,点击查看Google Cloud服务账号
  • 在Google Cloud控制台中,为该服务账号创建一个JSON格式的密钥,下载到本地后保存为dialogflow-agent-key.json(路径自己记好)

3. 设置认证环境变量

让SDK能自动识别你的密钥文件,在终端中执行:

export GOOGLE_APPLICATION_CREDENTIALS="/home/你的用户名/保存密钥的路径/dialogflow-agent-key.json"

如果想让这个配置永久生效,可以把这条命令添加到~/.bashrc文件末尾,然后执行source ~/.bashrc刷新配置。

4. 编写CLI版的Dialogflow调用脚本

你可以写一个类似pushtotalk.py的脚本,支持语音输入+意图响应的完整流程。下面是一个基础示例(需要额外安装语音处理依赖):

先安装语音相关依赖

sudo apt-get install portaudio19-dev python3-pyaudio
pip install speechrecognition pyttsx3

示例脚本 df_pushtotalk.py

import dialogflow_v2 as dialogflow
import os
import speech_recognition as sr
import pyttsx3

# 替换成你的Dialogflow Agent的Project ID
PROJECT_ID = "你的-dialogflow-project-id"
SESSION_ID = "cli-session-123"  # 自定义会话ID即可,用于区分不同对话

def detect_intent_audio(project_id, session_id, audio_file_path, language_code):
    session_client = dialogflow.SessionsClient()
    session = session_client.session_path(project_id, session_id)

    with open(audio_file_path, "rb") as audio_file:
        input_audio = audio_file.read()

    audio_config = dialogflow.types.InputAudioConfig(
        audio_encoding=dialogflow.enums.AudioEncoding.AUDIO_ENCODING_LINEAR_16,
        sample_rate_hertz=16000,
        language_code=language_code,
    )
    query_input = dialogflow.types.QueryInput(audio_config=audio_config)

    response = session_client.detect_intent(
        session=session, query_input=query_input
    )

    return response.query_result.fulfillment_text

def listen_and_respond():
    # 初始化语音识别和合成引擎
    r = sr.Recognizer()
    engine = pyttsx3.init()

    with sr.Microphone(sample_rate=16000) as source:
        print("请说话...")
        audio = r.listen(source)

    # 保存临时音频文件
    with open("temp_audio.wav", "wb") as f:
        f.write(audio.get_wav_data())

    # 调用Dialogflow意图并获取响应
    response_text = detect_intent_audio(PROJECT_ID, SESSION_ID, "temp_audio.wav", "zh-CN")
    print(f"Dialogflow响应: {response_text}")

    # 播放响应语音
    engine.say(response_text)
    engine.runAndWait()

if __name__ == "__main__":
    listen_and_respond()

运行脚本

python df_pushtotalk.py

运行后,对着麦克风说话,脚本会自动识别语音、调用Dialogflow意图,最后输出并播放响应内容,和pushtotalk.py的使用体验基本一致。

5. 简化版:文本直接调用意图

如果只是想快速测试文本触发意图,可以写一个更轻量化的脚本df_text_query.py:

import dialogflow_v2 as dialogflow
import sys

PROJECT_ID = "你的-dialogflow-project-id"
SESSION_ID = "cli-session-123"

def detect_intent_text(project_id, session_id, text, language_code):
    session_client = dialogflow.SessionsClient()
    session = session_client.session_path(project_id, session_id)

    text_input = dialogflow.types.TextInput(text=text, language_code=language_code)
    query_input = dialogflow.types.QueryInput(text=text_input)

    response = session_client.detect_intent(session=session, query_input=query_input)
    return response.query_result.fulfillment_text

if __name__ == "__main__":
    if len(sys.argv) < 2:
        print("用法: python df_text_query.py '你的查询文本'")
        sys.exit(1)
    query_text = sys.argv[1]
    response = detect_intent_text(PROJECT_ID, SESSION_ID, query_text, "zh-CN")
    print(f"Dialogflow响应: {response}")

运行方式:

python df_text_query.py "今天天气怎么样?"

注意事项

  • 确保你的Dialogflow Agent已经完成训练并发布,否则可能无法返回正确的意图响应
  • 如果遇到权限报错,检查服务账号是否拥有Dialogflow API的访问权限
  • Ubuntu 16.04默认的Python3版本是3.5,Dialogflow SDK支持该版本,若出现兼容性问题,可考虑使用虚拟环境隔离依赖

内容的提问来源于stack exchange,提问作者Jaime

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 08:23:59