使用Gradio Client调用Falcon 40B空间仅返回首个token问题排查
解决Gradio Client获取Falcon 40B结果仅返回首个Token的问题
问题出在目标Hugging Face Space的聊天接口采用流式输出模式,会分批次返回生成的token。直接调用response.result()只会获取流式返回的第一个片段,因此只拿到了"I"。
解决方案1:等待任务完成后获取完整结果
调用wait_for_completion()等待模型生成结束,再获取最终结果:
from gradio_client import Client instructions="""The following is a conversation between a highly knowledgeable and intelligent AI assistant, called Falcon, and a human user, called User. In the following interactions, User and Falcon will converse in natural language, and Falcon will answer User's questions. Falcon was built to be respectful, polite and inclusive. Falcon was built by the Technology Innovation Institute in Abu Dhabi. Falcon will never decline to answer a question, and always attempts to give an answer that User would be satisfied with. It knows a lot, and always tells the truth. The conversation begins.""" client = Client("https://huggingfaceh4-falcon-chat.hf.space/", serialize=False) response = client.submit( "Hey There! Respond to this with I like chicken.", [], instructions, 0.8, 0.6, fn_index=1 ) # 等待生成任务完成 response.wait_for_completion() # 获取完整结果 print(response.result())
解决方案2:遍历流式片段拼接完整回答
如果需要实时获取生成进度,可以遍历response.iterations()收集所有输出片段:
from gradio_client import Client instructions="""The following is a conversation between a highly knowledgeable and intelligent AI assistant, called Falcon, and a human user, called User. In the following interactions, User and Falcon will converse in natural language, and Falcon will answer User's questions. Falcon was built to be respectful, polite and inclusive. Falcon was built by the Technology Innovation Institute in Abu Dhabi. Falcon will never decline to answer a question, and always attempts to give an answer that User would be satisfied with. It knows a lot, and always tells the truth. The conversation begins.""" client = Client("https://huggingfaceh4-falcon-chat.hf.space/", serialize=False) response = client.submit( "Hey There! Respond to this with I like chicken.", [], instructions, 0.8, 0.6, fn_index=1 ) full_answer = "" # 遍历所有流式输出片段 for chunk in response.iterations(): # 流式输出通常返回累积生成内容,直接覆盖即可 full_answer = chunk[0][1] # 输出完整对话结果 print([["Hey There! Respond to this with I like chicken.", full_answer]])
方案1更简洁,适合只需要最终结果的场景;方案2适合需要实时展示生成过程的场景,两种方式都能获取完整回答。
内容的提问来源于stack exchange,提问作者Yupery
相关产品推荐
相关产品推荐

