You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

LangChain简历解析遇ValueError:缺失输入键'\n "ContactInformation"'

解决LangChain+Pydantic简历解析工具的ValueError问题

错误原因

报错ValueError: Missing some input keys: {'\n "ContactInformation"'}的核心原因:

  • LangChain的PromptTemplate默认将{}视为变量占位符,你硬写在PROMPT_2里的JSON结构包含大量普通大括号,被模板引擎误识别为未定义的输入变量。
  • 代码中已定义PydanticOutputParser但未在Prompt中使用其生成的格式规则,反而手动编写JSON schema,既冗余又引发解析冲突。

修复方案

1. 修正Prompt模板

移除手动编写的JSON schema,改用PydanticOutputParser生成的format_instructions,确保输出格式与Pydantic模型完全一致,同时避免大括号误解析问题。

PROMPT_2 = """你将收到一份简历内容:```{resume}```

请根据简历提取个人信息,严格遵循以下格式要求返回结果:
{format_instructions}

注意:如果某个字段没有对应信息,必须设为null;邮箱需包含@符号;联系电话为纯数字字符串;姓名通常在简历开头无标签位置;社交链接需为有效URL。"""

2. 修正核心函数

确保Prompt模板中包含{format_instructions}占位符,同时保留partial_variables传递解析规则:

import json
from langchain.chat_models import ChatOpenAI
from langchain.prompts import PromptTemplate
from langchain.chains import LLMChain
from langchain.output_parsers import PydanticOutputParser
from pydantic import BaseModel, Optional, List, Union

def ExtractInformationFromResume(resume: str) -> OutputFormat:
    llm = ChatOpenAI(
        openai_api_key='*********************************************',
        temperature=0.1,  # 降低随机性,提升格式稳定性
        model_name="gpt-3.5-turbo"
    )

    # 定义输出解析器
    parser = PydanticOutputParser(pydantic_object=OutputFormat)

    # 创建带格式说明的Prompt模板
    prompt_template = PromptTemplate(
        input_variables=["resume"],
        template=PROMPT_2,
        partial_variables={"format_instructions": parser.get_format_instructions()},
    )

    # 构建LLM链
    chain = LLMChain(llm=llm, prompt=prompt_template)

    # 执行解析
    print(f"Resume Length: {len(resume)} characters")
    response = chain.run({"resume": resume})
    
    print("Raw LLM Response:", response)

    # 清理并解析响应
    try:
        # 移除可能的Markdown标记
        cleaned_response = response.strip().strip("```json").strip("```")
        return OutputFormat(**json.loads(cleaned_response))
    except json.JSONDecodeError as e:
        raise ValueError(f"解析响应失败:{str(e)},原始响应:{response}")

3. 优化Pydantic模型(可选)

给字段添加描述,让LLM更清晰理解字段要求,提升解析准确性:

from pydantic import BaseModel, Field, Optional, List, Union

class ContactInformation(BaseModel):
    Name: Optional[str] = Field(None, description="个人姓名,通常在简历开头无标签位置")
    Email: Optional[str] = Field(None, description="包含@符号的邮箱地址")
    Contact: Optional[str] = Field(None, description="纯数字字符串格式的联系电话")
    Links: Optional[List[str]] = Field(None, description="社交平台或个人主页的有效URL数组")

class Experience(BaseModel):
    title: Optional[str] = Field(None, description="职位名称")
    company: Optional[str] = Field(None, description="公司名称")
    duration: Optional[str] = Field(None, description="任职时长")

# 其余模型类同理添加Field描述

额外注意事项

  • 降低temperature值(如0.1)可减少LLM输出的随机性,避免返回不符合格式的内容。
  • 对LLM返回的响应做清理,去除可能的Markdown标记(如json),避免JSON解析失败。
  • 确保Pydantic模型的字段名与Prompt要求完全一致,避免大小写或命名差异导致解析错误。

内容的提问来源于stack exchange,提问作者sarthak basnet

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 20:59:53