基于HuggingFace与LangChain的Few-Shot简历解析问题求助
问题解决与适配方案
一、'ContactInformation'解析错误修复
Flan-T5 Base属于小参数模型,在Few-Shot提示下对结构化输出的约束理解不足,容易出现只输出字段名、格式混乱或字段缺失的情况,直接导致你遇到的解析错误。以下是具体修复步骤:
强化Few-Shot示例的结构化约束
把示例里的ContactInformation字段结构写死,明确子字段(电话、邮箱、地址等)的格式,用JSON作为强制输出规范。比如:示例简历: 张三 电话:138xxxx1234 邮箱:zhangsan@xxx.com 地址:北京市朝阳区 解析结果: { "ContactInformation": { "phone": "138xxxx1234", "email": "zhangsan@xxx.com", "address": "北京市朝阳区" }, "Education": {...} }同时在提示词末尾加硬约束:必须输出完整JSON结构,不得仅输出字段名,确保每个字段都有对应内容。
添加格式校验与重试逻辑
在代码里加入JSON解析校验,一旦出现字段缺失或格式错误,自动调整提示词重试。示例代码片段:import json from langchain.llms import HuggingFacePipeline from transformers import AutoTokenizer, AutoModelForSeq2SeqLM, pipeline def parse_resume(resume_text, llm, few_shot_examples): base_prompt = f"{few_shot_examples}\n请解析以下简历,输出JSON格式结果:\n{resume_text}" result = llm(base_prompt) try: parsed_data = json.loads(result) if not parsed_data.get("ContactInformation"): raise ValueError("ContactInformation字段为空或缺失") return parsed_data except (json.JSONDecodeError, ValueError): # 重试时强化格式要求 retry_prompt = f"{few_shot_examples}\n请严格按照示例的JSON格式解析,必须包含完整的ContactInformation字段:\n{resume_text}" retry_result = llm(retry_prompt) return json.loads(retry_result)调整模型生成参数
降低模型随机性,强制贴合示例格式:model_name = "google/flan-t5-base" tokenizer = AutoTokenizer.from_pretrained(model_name) model = AutoModelForSeq2SeqLM.from_pretrained(model_name) pipe = pipeline( "text2text-generation", model=model, tokenizer=tokenizer, temperature=0.2, # 降低随机性 max_new_tokens=512, # 避免截断输出 do_sample=False # 生成确定性结果 ) llm = HuggingFacePipeline(pipeline=pipe)
二、低配置笔记本的模型替代方案
如果Flan-T5 Base仍占用资源过高,试试这些更轻量化的方案:
改用更小的T5衍生模型
直接用google/flan-t5-small(60M参数)或google/flan-t5-tiny(30M参数),推理速度提升数倍,简历解析这类结构化任务在Few-Shot加持下,基本精度不会受太大影响。模型量化压缩
用bitsandbytes库做4-bit/8-bit量化,显存占用直接砍半甚至更多,几乎不损失精度:from transformers import AutoModelForSeq2SeqLM, BitsAndBytesConfig bnb_config = BitsAndBytesConfig( load_in_4bit=True, bnb_4bit_use_double_quant=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=torch.bfloat16 ) model = AutoModelForSeq2SeqLM.from_pretrained( "google/flan-t5-base", quantization_config=bnb_config, device_map="auto" )规则+轻量NER模型混合方案
用正则提取邮箱、手机号这类格式固定的信息,再用轻量NER模型(如dslim/bert-base-NER)提取姓名、地址,混合方案速度快、准确率高:import re from transformers import pipeline ner_pipeline = pipeline("ner", model="dslim/bert-base-NER") def extract_contact_info(resume_text): # 正则提取邮箱、手机号 email = re.findall(r'\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Z|a-z]{2,}\b', resume_text) phone = re.findall(r'\b1[3-9]\d{9}\b', resume_text) # NER提取姓名、地址 ner_results = ner_pipeline(resume_text) name = ' '.join([item['word'] for item in ner_results if item['entity'] == 'B-PER']) address = ' '.join([item['word'] for item in ner_results if item['entity'] == 'B-LOC']) return { "ContactInformation": { "name": name, "phone": phone[0] if phone else "", "email": email[0] if email else "", "address": address } }CPU推理优化
开启CPU多线程或转换为ONNX格式,提升CPU推理速度:# 多线程优化 import torch torch.set_num_threads(4) # 按CPU核心数调整 # ONNX格式转换 from optimum.onnxruntime import ORTModelForSeq2SeqLM model = ORTModelForSeq2SeqLM.from_pretrained("google/flan-t5-base", export=True)
内容的提问来源于stack exchange,提问作者sarthak basnet
相关产品推荐
相关产品推荐

