You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用LangChain+Bloom模型生成会议纪要时遇ValidationError问题求助

问题原因

你直接将transformers加载的AutoModelForCausalLM实例传给load_summarize_chain的model参数,而LangChain的摘要链要求传入LangChain封装的LLM类实例(比如HuggingFacePipeline),而非原生transformers模型对象,因此触发类型验证错误。

解决方案

用LangChain的HuggingFacePipeline类,将transformers加载的模型和tokenizer封装成符合LangChain规范的LLM实例,再传入摘要链即可解决问题。

修改后的完整代码

from transformers import AutoModelForCausalLM, AutoTokenizer, pipeline
from langchain.chains.summarize import load_summarize_chain
from langchain.docstore.document import Document
from langchain.prompts import PromptTemplate
from langchain.text_splitter import CharacterTextSplitter
from langchain.llms import HuggingFacePipeline

checkpoint = "bigscience/bloom-560m"
tokenizer = AutoTokenizer.from_pretrained(checkpoint)
model = AutoModelForCausalLM.from_pretrained(checkpoint)

# 用transformers pipeline封装模型,配置生成参数
text_generation_pipeline = pipeline(
    "text-generation",
    model=model,
    tokenizer=tokenizer,
    max_new_tokens=500,  # 匹配设定的target_len
    temperature=0.3,  # 降低随机性,保证生成内容严谨
    top_p=0.9,
    repetition_penalty=1.1
)

# 转换为LangChain兼容的LLM实例
llm = HuggingFacePipeline(pipeline=text_generation_pipeline)

transcript_file = "/content/transcript/transcript.txt" 
with open(transcript_file, encoding='latin-1') as file:
    documents = file.read()

text_splitter = CharacterTextSplitter(
    chunk_size=3000,
    chunk_overlap=200,
    length_function=len
)
texts = text_splitter.split_text(documents)
docs = [Document(page_content=t) for t in texts[:]]

target_len = 500
prompt_template = """Act as a professional technical meeting minutes writer. 
Tone: formal
Format: Technical meeting summary
Tasks:
- Highlight action items and owners
- Highlight the agreements
- Use bullet points if needed

{text}

CONCISE SUMMARY IN ENGLISH:"""
PROMPT = PromptTemplate(template=prompt_template, input_variables=["text"])
refine_template = (
    "Your job is to produce a final summary\n"
    "We have provided an existing summary up to a certain point: {existing_answer}\n"
    "We have the opportunity to refine the existing summary"
    "(only if needed) with some more context below.\n"
    "------------\n"
    "{text}\n"
    "------------\n"
    f"Given the new context, refine the original summary in English within {target_len} words: following the format\n"
    "Participants: <participants>\n"
    "Discussed: <Discussed-items>\n"
    "Follow-up actions: <a-list-of-follow-up-actions-with-owner-names>\n"
    "If the context isn't useful, return the original summary. Highlight agreements and follow-up actions and owners."
)
refine_prompt = PromptTemplate(
    input_variables=["existing_answer", "text"],
    template=refine_template,
)

chain = load_summarize_chain(
    llm=llm,  # 传入封装后的LLM实例
    chain_type="refine",
    return_intermediate_steps=True,
    question_prompt=PROMPT,
    refine_prompt=refine_prompt
)
result = chain({"input_documents": docs}, return_only_outputs=True)

# 输出最终会议纪要
print(result['output_text'])

关键修改说明

  1. 新增HuggingFacePipeline导入,用transformers的text-generation pipeline包装模型,同时配置生成参数(比如max_new_tokens控制输出长度)
  2. 将封装后的llm实例传入load_summarize_chain的llm参数(替代原代码的model参数,更符合LangChain规范)
  3. 修正了refine模板中的换行问题,确保格式指令能正确生效

内容的提问来源于stack exchange,提问作者Shaik Naveed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 21:20:35