LangChain简历解析遇ValueError:缺失输入键'\n "ContactInformation"'
解决LangChain+Pydantic简历解析工具的ValueError问题
错误原因
报错ValueError: Missing some input keys: {'\n "ContactInformation"'}的核心原因:
- LangChain的
PromptTemplate默认将{}视为变量占位符,你硬写在PROMPT_2里的JSON结构包含大量普通大括号,被模板引擎误识别为未定义的输入变量。 - 代码中已定义
PydanticOutputParser但未在Prompt中使用其生成的格式规则,反而手动编写JSON schema,既冗余又引发解析冲突。
修复方案
1. 修正Prompt模板
移除手动编写的JSON schema,改用PydanticOutputParser生成的format_instructions,确保输出格式与Pydantic模型完全一致,同时避免大括号误解析问题。
PROMPT_2 = """你将收到一份简历内容:```{resume}``` 请根据简历提取个人信息,严格遵循以下格式要求返回结果: {format_instructions} 注意:如果某个字段没有对应信息,必须设为null;邮箱需包含@符号;联系电话为纯数字字符串;姓名通常在简历开头无标签位置;社交链接需为有效URL。"""
2. 修正核心函数
确保Prompt模板中包含{format_instructions}占位符,同时保留partial_variables传递解析规则:
import json from langchain.chat_models import ChatOpenAI from langchain.prompts import PromptTemplate from langchain.chains import LLMChain from langchain.output_parsers import PydanticOutputParser from pydantic import BaseModel, Optional, List, Union def ExtractInformationFromResume(resume: str) -> OutputFormat: llm = ChatOpenAI( openai_api_key='*********************************************', temperature=0.1, # 降低随机性,提升格式稳定性 model_name="gpt-3.5-turbo" ) # 定义输出解析器 parser = PydanticOutputParser(pydantic_object=OutputFormat) # 创建带格式说明的Prompt模板 prompt_template = PromptTemplate( input_variables=["resume"], template=PROMPT_2, partial_variables={"format_instructions": parser.get_format_instructions()}, ) # 构建LLM链 chain = LLMChain(llm=llm, prompt=prompt_template) # 执行解析 print(f"Resume Length: {len(resume)} characters") response = chain.run({"resume": resume}) print("Raw LLM Response:", response) # 清理并解析响应 try: # 移除可能的Markdown标记 cleaned_response = response.strip().strip("```json").strip("```") return OutputFormat(**json.loads(cleaned_response)) except json.JSONDecodeError as e: raise ValueError(f"解析响应失败:{str(e)},原始响应:{response}")
3. 优化Pydantic模型(可选)
给字段添加描述,让LLM更清晰理解字段要求,提升解析准确性:
from pydantic import BaseModel, Field, Optional, List, Union class ContactInformation(BaseModel): Name: Optional[str] = Field(None, description="个人姓名,通常在简历开头无标签位置") Email: Optional[str] = Field(None, description="包含@符号的邮箱地址") Contact: Optional[str] = Field(None, description="纯数字字符串格式的联系电话") Links: Optional[List[str]] = Field(None, description="社交平台或个人主页的有效URL数组") class Experience(BaseModel): title: Optional[str] = Field(None, description="职位名称") company: Optional[str] = Field(None, description="公司名称") duration: Optional[str] = Field(None, description="任职时长") # 其余模型类同理添加Field描述
额外注意事项
- 降低
temperature值(如0.1)可减少LLM输出的随机性,避免返回不符合格式的内容。 - 对LLM返回的响应做清理,去除可能的Markdown标记(如
json),避免JSON解析失败。 - 确保Pydantic模型的字段名与Prompt要求完全一致,避免大小写或命名差异导致解析错误。
内容的提问来源于stack exchange,提问作者sarthak basnet
相关产品推荐
相关产品推荐

