使用LangChain+Bloom模型生成会议纪要时遇ValidationError问题求助
问题原因
你直接将transformers加载的AutoModelForCausalLM实例传给load_summarize_chain的model参数,而LangChain的摘要链要求传入LangChain封装的LLM类实例(比如HuggingFacePipeline),而非原生transformers模型对象,因此触发类型验证错误。
解决方案
用LangChain的HuggingFacePipeline类,将transformers加载的模型和tokenizer封装成符合LangChain规范的LLM实例,再传入摘要链即可解决问题。
修改后的完整代码
from transformers import AutoModelForCausalLM, AutoTokenizer, pipeline from langchain.chains.summarize import load_summarize_chain from langchain.docstore.document import Document from langchain.prompts import PromptTemplate from langchain.text_splitter import CharacterTextSplitter from langchain.llms import HuggingFacePipeline checkpoint = "bigscience/bloom-560m" tokenizer = AutoTokenizer.from_pretrained(checkpoint) model = AutoModelForCausalLM.from_pretrained(checkpoint) # 用transformers pipeline封装模型,配置生成参数 text_generation_pipeline = pipeline( "text-generation", model=model, tokenizer=tokenizer, max_new_tokens=500, # 匹配设定的target_len temperature=0.3, # 降低随机性,保证生成内容严谨 top_p=0.9, repetition_penalty=1.1 ) # 转换为LangChain兼容的LLM实例 llm = HuggingFacePipeline(pipeline=text_generation_pipeline) transcript_file = "/content/transcript/transcript.txt" with open(transcript_file, encoding='latin-1') as file: documents = file.read() text_splitter = CharacterTextSplitter( chunk_size=3000, chunk_overlap=200, length_function=len ) texts = text_splitter.split_text(documents) docs = [Document(page_content=t) for t in texts[:]] target_len = 500 prompt_template = """Act as a professional technical meeting minutes writer. Tone: formal Format: Technical meeting summary Tasks: - Highlight action items and owners - Highlight the agreements - Use bullet points if needed {text} CONCISE SUMMARY IN ENGLISH:""" PROMPT = PromptTemplate(template=prompt_template, input_variables=["text"]) refine_template = ( "Your job is to produce a final summary\n" "We have provided an existing summary up to a certain point: {existing_answer}\n" "We have the opportunity to refine the existing summary" "(only if needed) with some more context below.\n" "------------\n" "{text}\n" "------------\n" f"Given the new context, refine the original summary in English within {target_len} words: following the format\n" "Participants: <participants>\n" "Discussed: <Discussed-items>\n" "Follow-up actions: <a-list-of-follow-up-actions-with-owner-names>\n" "If the context isn't useful, return the original summary. Highlight agreements and follow-up actions and owners." ) refine_prompt = PromptTemplate( input_variables=["existing_answer", "text"], template=refine_template, ) chain = load_summarize_chain( llm=llm, # 传入封装后的LLM实例 chain_type="refine", return_intermediate_steps=True, question_prompt=PROMPT, refine_prompt=refine_prompt ) result = chain({"input_documents": docs}, return_only_outputs=True) # 输出最终会议纪要 print(result['output_text'])
关键修改说明
- 新增
HuggingFacePipeline导入,用transformers的text-generationpipeline包装模型,同时配置生成参数(比如max_new_tokens控制输出长度) - 将封装后的
llm实例传入load_summarize_chain的llm参数(替代原代码的model参数,更符合LangChain规范) - 修正了refine模板中的换行问题,确保格式指令能正确生效
内容的提问来源于stack exchange,提问作者Shaik Naveed
相关产品推荐
相关产品推荐

