LangChain调用Cohere Embeddings报类型错误:texts应为字符串
问题分析与解决:CohereEmbeddings调用时的类型错误
错误日志
Retrying langchain_cohere.embeddings.CohereEmbeddings.embed_with_retry.<locals>._embed_with_retry in 4.0 seconds as it raised UnprocessableEntityError: status_code: 422, body: {'message': 'invalid type: parameter texts is of type object but should be of type string. For proper usage, please refer to https://docs.cohere.com/v1/reference/embed'}
相关代码
用户交互逻辑
question = input("Ask your question: ") chat_history.append(f"user: {question}") print("********************************************") print(chat_history, type(chat_history)) print("********************************************") while question != "Bye": chat_history_str = "\n".join(chat_history) print (chat_history_str, type(chat_history_str)) print("++++++++++++++++++++++++++++++++++++++++++++") response = rag_chain.invoke( { 'question': question, 'chat_history': chat_history_str } ) print(response) print("----------------------------------------------") chat_history.append(f"AI: {response}")
RAG链定义
rag_chain = ( {"context": retriever | format_docs, "question": RunnablePassthrough().pick("question"), "input": RunnablePassthrough().pick("question"), "chat_history": RunnablePassthrough().pick("chat_history")} | final_prompt | llm | StrOutputParser() )
排查过程
- 确认
chat_history_str为字符串类型,排除该参数的类型问题 - 将
retriever替换为RunnablePassthrough()后,链可正常运行(无上下文) - 单独调用
retriever.invoke()返回结果正常,锁定问题出在context获取阶段(即retriever | format_docs环节)
问题定位
错误中提到的"parameter texts is of type object",指的是传入CohereEmbeddings的文本参数是文档对象而非字符串。根源在于format_docs函数没有正确将检索到的文档列表转换为纯字符串。
Retriever返回的是文档对象列表(如Qdrant的Document实例),如果format_docs没有提取文档的文本内容并拼接成字符串,而是直接返回了文档对象或未处理的列表,就会导致Cohere的嵌入接口接收到非字符串类型的参数,触发422错误。
修复方案
确保format_docs函数返回单个字符串,示例如下:
def format_docs(docs): # 提取每个文档的page_content并拼接 return "\n\n".join(doc.page_content for doc in docs)
内容的提问来源于stack exchange,提问作者Akshitha Rao
相关产品推荐
相关产品推荐

