You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Raspberry Pi 4与AWS Lex的语音聊天Bot性能优化问询

问题背景

我正在开发一款基于Raspberry Pi 4和AWS Lex的语音聊天机器人,目前面临两个核心问题:

  1. 大量判断逻辑与检测操作导致录音模块启动延迟,用户说话时常无法进入录音阶段。
  2. 使用subprocess调用Raspberry Pi专属GPU加速播放器omxplayer播放状态视频时,无法直接通过subprocess终止进程,仅os.killpg方法有效,但相关进程检测与终止逻辑可能拖慢主流程效率。

目前已用两个线程分别处理热词检测与运动检测,考虑过将进程检测放入独立线程,但担心增加代码复杂度;尝试过多进程优化但未见效,也已简化过判断逻辑。

核心代码片段

进程终止逻辑

if loadingVid.poll() is None:           
    os.killpg(os.getpgid(loadingVid.pid), signal.SIGTERM)

录音函数

def record():  
    """Record audio from the microphone"""  
    os.system("sox -d -t wavpcm -c 1 -b 16 -r 16000 -e signed-integer --endian little - silence 1 0 1% 5 0.8t 5% -highpass 300> request.wav")

主函数

def main():
    """
    Main function:
    1. Load environment variables and start idle video.
    2. Initialize Amazon Lex runtime client.
    3. Set max waiting time, current session id, and last response to initial values.
    4. Enter an infinite loop:
        5. If the last response's session state's intent state is "Fulfilled" or "Failed":
            a. Wait for a hot word and return the time elapsed waiting for it.
            b. If the idle time duration is greater than the max waiting time, start a new session and play a greeting video. Otherwise, play a confirmation video.
        6. If the last response is None:
            a. Wait for a hot word and return the time elapsed waiting for it.
            b. Start a new session and play a greeting video.
        7. Stop the listening video and start a loading video.
        8. Record audio and send it to the Amazon Lex runtime to get a response.
        9. Handle the response by playing an audio file, displaying an image, or playing a video.
        10. Stop the loading video and start the idle video again.
    """

    dotenv.load_dotenv()
    hologram.minimze()
    hologram.hide_cursor()
    # hologram.bench_idle()
    idle = hologram.play_idle("/home/alexa/Videos/menu.mp4", 5)
    lexruntimev2 = boto3.client(
        "lexv2-runtime",
        aws_access_key_id=os.environ.get("aws_access_key_id"),
        aws_secret_access_key=os.environ.get("aws_secret_access_key"),
        region_name="us-east-1",
    )

    maxWaitingTime = 30.0
    currentSessionId = None
    last_response = None
    current_process = None
    listeningVid = None
    intermediate_vid = False
    while True:
        if last_response and (
            last_response["sessionState"]["intent"]["state"] == "Fulfilled"
            or last_response["sessionState"]["intent"]["state"] == "Failed"
        ):
        
            if idle.poll() is not None:
                idle = hologram.play_idle("/home/alexa/Videos/menu.mp4", 5)
            # wait for hot word and return time elased waiting for it
            idleTimeDuration = triggers.wait_for_triggers()
            if idleTimeDuration > maxWaitingTime:
                currentSessionId = lex.newSession()
                greeting_vid = lex.say_greeting()
            else:
                listeningVid = lex.say_confirm_listening()
        elif last_response == None:
            triggers.wait_for_triggers()
            currentSessionId = lex.newSession()
            lex.say_greeting()
        # if this is a follow-up question (slot elicitation)
        else:
            listeningVid= hologram.play_idle("/home/alexa/project/video/speaking.mp4",8) 
        lex.record()
        
        # say one moment please
        loadingVid = lex.say_one_moment()

        # if the idle vid is still running (return value is still none)
        if idle.poll() is None:
            # then kill it
            os.killpg(os.getpgid(idle.pid), signal.SIGTERM)
        
        if listeningVid is not None:
            # then kill the listening
            if listeningVid.poll() is None:
                os.killpg(os.getpgid(listeningVid.pid), signal.SIGTERM)
        response, responseAssetURL = lex.recognize_audio(lexruntimev2, currentSessionId)
        last_response = response
        video_case_handling = True
        Image_case_handling = True

        if responseAssetURL != None:
            if ".png" not in responseAssetURL.lower():
                audio.play_audio("audio/lex_response.mpeg")
                video_case_handling = False
            if ".png" in responseAssetURL.lower():
                Image_case_handling = False
                feh = hologram.display_image(
                    responseAssetURL, "/home/alexa/project/images/image.png"
                )

            if Image_case_handling:
                hologram.play_with_omx(responseAssetURL, 9)

        # if the loading vid is still running (return value is still none)
        if loadingVid.poll() is None:
            # then kill it
            os.killpg(os.getpgid(loadingVid.pid), signal.SIGTERM)
        
        if responseAssetURL == None:
            intermediate_vid = True
            intermediate = hologram.play_idle(
                "/home/alexa/project/video/speaking.mp4", 8
            )
        if video_case_handling:
            audio.play_audio("audio/lex_response.mpeg")
            if intermediate_vid:
                os.killpg(os.getpgid(intermediate.pid), signal.SIGTERM)
                intermediate_vid = False

        if Image_case_handling == False:
            idle = hologram.play_idle("/home/alexa/Videos/menu.mp4", 5) 
            os.system("pkill feh")
优化思路方向

录音延迟优化

  • 梳理主流程中lex.record()执行前的所有操作,拆分出可并行或预执行的步骤,比如将非必要的视频状态更新、进程检测逻辑延后到录音启动后处理。
  • 对sox录音命令做参数优化,调整silence检测阈值、预处理流程,减少启动阶段的耗时。
  • 提前初始化录音模块,比如在热词触发后立即预加载录音相关资源,而非等待所有判断逻辑完成再启动。

进程管理优化

  • 封装专门的进程管理器类,集中维护所有视频播放进程的启动、检测、终止逻辑,避免在主流程中分散大量重复判断。
  • 尝试异步触发进程终止操作,发送终止信号后不同步等待结果,减少主流程阻塞时间。
  • 探索omxplayer的原生终止方式,比如通过stdin发送q命令,替代os.killpg,降低进程组管理的复杂度。
可改进点
  • 重构主函数中的条件判断,合并重复逻辑(如多次出现的热词等待、会话初始化代码),减少代码冗余与判断耗时。
  • 将录音函数中的os.system替换为subprocess模块调用,更好地控制录音进程,同时避免shell层的额外开销。
  • 用统一的数据结构(如字典)管理所有视频播放进程对象,避免零散变量带来的判断混乱。
注意事项
  • 新增线程/进程时需注意线程安全,对共享变量(如last_response、进程对象)的访问要加锁,避免竞态条件。
  • Raspberry Pi资源有限,并行线程/进程数量不宜过多,防止CPU或内存占用过高导致整体性能下降。
  • 修改进程终止逻辑时,需测试进程已正常退出的极端情况,添加异常捕获避免os.killpg调用失败引发的程序崩溃。
  • 录音延迟问题需先做性能分析,明确是判断逻辑耗时还是录音命令本身启动慢,再针对性优化,避免盲目调整。

内容的提问来源于stack exchange,提问作者Ali Redha

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 22:40:57