You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python TTS机器人中如何让ChatGPT在完整语句处停止响应?

解决ChatGPT回复截断及上下文补全异常问题

问题情况

我是编程新手,首个项目是基于Python的语音交互ChatGPT TTS机器人,已实现语音提问与TTS回复功能,但遇到两个核心问题:

  • 调整max_tokens(设为40或100)后,ChatGPT的回复经常语句未完成就被截断,例如:ChatGPT:‘草是绿色的,已经有一段时间没修剪了,还’
  • 后续提问时,ChatGPT会补全上轮未说完的内容,例如用户问“你了解奶牛吗?”,回复却是‘有一些粉色的花。是的,我可以告诉你一些很棒的奶牛相关内容’
    尝试过分句函数和调整token数量都无效,希望通过代码实现让ChatGPT在完整语句结束时停止响应,等待下一个问题。

原代码如下:

# Sentence chunking function
def chunk_text_into_sentences(text):
    sentences = text.split('. ')
    return sentences   

# Main Mic, Chat History
def main():
    recording = False
    paused = False

    while True:
        if keyboard.is_pressed("m"):
            recording = True
            print("Microphone ON. Say your question...")
            while recording:
                if keyboard.is_pressed("p"):  # Pause
                    paused = True
                    print("Bot paused.")
                    while keyboard.is_pressed("p"):  # Wait for 'p' key to be released
                        time.sleep(0.1)
                if not paused:
                    with sr.Microphone() as source:
                        recognizer = sr.Recognizer()
                        audio = recognizer.listen(source)
                        try:
                            transcription = recognizer.recognize_google(audio)
                            print("Transcription:", transcription) 
                            if transcription.lower() == "n":
                                recording = False
                                print("Microphone OFF.")
                            else:
                                text = transcription
                                if text:
                                    print("You said:", text)
                                    response = generate_response(text) 
                                    print("Response:", response)
                                    text_to_speech_elevenlabs(response)
                                    chat_history.append({"role": "user", "content": text})
                                    chat_history.append({"role": "assistant", "content": response})
                        except sr.UnknownValueError:
                            pass

                if keyboard.is_pressed("u"):
                    paused = False
                    print("Bot resumed.")
                    while keyboard.is_pressed("u"):  # Wait for 'u' key to be released
                        time.sleep(0.1)

def generate_response(prompt, max_tokens=40):
    response = openai.ChatCompletion.create(
        model="gpt-3.5-turbo-16k-0613",
        messages=chat_history,
        max_tokens=max_tokens,
        n=1,
        stop=None,
        temperature=1,
    )
    return response.choices[0].message['content']

问题根源

  1. 固定max_tokens适配性差:固定token数无法匹配不同语句的长度,短语句浪费token额度,长语句直接被截断。
  2. 不完整回复混入聊天历史:将截断的不完整回复加入chat_history,导致AI在下一轮对话中自动补全上轮未完成的内容。
  3. 分句函数未适配中文:仅用英文句号加空格分割,无法识别中文语句结束标点(。),起不到完整语句筛选作用。

解决方案

1. 利用stop参数控制AI终止生成

在调用OpenAI接口时,设置stop参数为中文语句结束标点,让AI在生成到完整语句时自动停止,避免截断。

2. 过滤不完整回复

检查AI回复是否以结束标点结尾,若未结尾则不加入聊天历史,避免影响后续对话。

3. 优化分句逻辑

适配中文标点,正确拆分完整语句。

4. 调整max_tokens为合理上限

设置足够大的max_tokens(如200),给AI足够空间生成完整语句,再通过stop参数控制终止。

修改后的完整代码

import keyboard
import time
import speech_recognition as sr
import openai

# 全局聊天历史
chat_history = []

# 适配中文的分句函数
def chunk_text_into_sentences(text):
    # 中文语句结束标点集合
    end_punctuations = {'。', '!', '?', '.', '!', '?'}
    sentences = []
    current_sentence = []
    for char in text:
        current_sentence.append(char)
        if char in end_punctuations:
            sentences.append(''.join(current_sentence))
            current_sentence = []
    # 处理最后一段未完成的语句
    if current_sentence:
        sentences.append(''.join(current_sentence))
    return sentences

def text_to_speech_elevenlabs(text):
    # 保留原TTS逻辑,此处省略具体实现
    pass

def generate_response(prompt, max_tokens=200):
    # 更新聊天历史,加入当前用户提问
    chat_history.append({"role": "user", "content": prompt})
    
    response = openai.ChatCompletion.create(
        model="gpt-3.5-turbo-16k-0613",
        messages=chat_history,
        max_tokens=max_tokens,
        n=1,
        # 设置中文语句结束标点为停止符
        stop=['。', '!', '?'],
        temperature=0.7,  # 降低随机性,让回复更稳定
    )
    
    assistant_response = response.choices[0].message['content']
    # 移除刚加入的用户提问,后续统一处理完整回复
    chat_history.pop()
    return assistant_response

# Main Mic, Chat History
def main():
    recording = False
    paused = False

    while True:
        if keyboard.is_pressed("m"):
            recording = True
            print("Microphone ON. Say your question...")
            while recording:
                if keyboard.is_pressed("p"):  # Pause
                    paused = True
                    print("Bot paused.")
                    while keyboard.is_pressed("p"):  # Wait for 'p' key to be released
                        time.sleep(0.1)
                if not paused:
                    with sr.Microphone() as source:
                        recognizer = sr.Recognizer()
                        audio = recognizer.listen(source)
                        try:
                            transcription = recognizer.recognize_google(audio)
                            print("Transcription:", transcription) 
                            if transcription.lower() == "n":
                                recording = False
                                print("Microphone OFF.")
                            else:
                                text = transcription
                                if text:
                                    print("You said:", text)
                                    response = generate_response(text) 
                                    print("Response:", response)
                                    
                                    # 检查回复是否为完整语句
                                    end_punctuations = {'。', '!', '?', '.', '!', '?'}
                                    is_complete = response.strip()[-1] in end_punctuations if response.strip() else False
                                    
                                    if is_complete:
                                        # 仅将完整回复加入聊天历史
                                        text_to_speech_elevenlabs(response)
                                        chat_history.append({"role": "user", "content": text})
                                        chat_history.append({"role": "assistant", "content": response})
                                    else:
                                        print("Warning: 回复不完整,未加入聊天历史")
                                        # 可选:可以再次调用接口补全,但需注意token消耗
                        except sr.UnknownValueError:
                            print("无法识别语音,请重新输入")

                if keyboard.is_pressed("u"):
                    paused = False
                    print("Bot resumed.")
                    while keyboard.is_pressed("u"):  # Wait for 'u' key to be released
                        time.sleep(0.1)

if __name__ == "__main__":
    main()

关键修改说明

  • generate_response函数:加入stop=['。', '!', '?']参数,让AI遇到中文结束标点时停止生成;调整temperature为0.7降低随机性;临时加入用户提问后再移除,等确认回复完整后统一加入聊天历史。
  • 完整回复校验:在主函数中检查回复结尾是否为结束标点,仅将完整回复加入聊天历史,避免不完整内容干扰后续对话。
  • 分句函数优化:适配中文标点,正确拆分完整语句。
  • max_tokens调整:设置为200,给AI足够生成完整语句的空间,同时通过stop参数避免不必要的长回复。

内容的提问来源于stack exchange,提问作者Jessica

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 15:29:58