使用本地Llama2-Chat-13B时PandasQueryEngine执行代码报错求助
问题解决思路与替代工具推荐
问题原因分析
报错核心是本地Llama2-Chat-13B生成的输出包含自然语言说明文本,而非纯可执行Python代码,导致ast.parse()无法解析。这并非未使用OpenAI的问题,而是开源模型对PandasQueryEngine的prompt指令遵循度不如闭源模型,额外生成了友好性说明文字(如"Sure, I'd be happy to help!..."),这类非代码内容触发了缩进错误。
解决思路
1. 自定义输出处理器
重写output_processor,过滤非代码内容,仅提取可执行Python代码:
def custom_output_processor(output: str) -> str: # 提取```包裹的代码块 if "```python" in output: code_block = output.split("```python")[1].split("```")[0].strip() return code_block # 清理自然语言前缀,保留代码行 lines = output.split("\n") code_lines = [line for line in lines if line.strip().startswith(("df.", "pd.", "print"))] return "\n".join(code_lines) # 初始化查询引擎时指定自定义处理器 query_engine = PandasQueryEngine( df=df, verbose=True, service_context=service_context, output_processor=custom_output_processor )
2. 优化prompt模板
修改系统prompt,强制模型仅输出纯Python代码:
from llama_index.prompts.prompts import PandasPrompt CUSTOM_PANDAS_PROMPT = PandasPrompt( prompt="""给定pandas dataframe `df`,编写Python代码片段回答用户问题。 仅输出Python代码,无解释、无额外文本。 问题:{query_str} 代码:""" ) query_engine = PandasQueryEngine( df=df, verbose=True, service_context=service_context, pandas_prompt=CUSTOM_PANDAS_PROMPT )
3. 调整模型参数
将模型temperature设为0.1或更低,降低输出随机性;同时确保使用Llama2-Chat专属的[INST]...[/INST]聊天模板,提升指令遵循度。
替代工具推荐
1. LangChain Pandas Agent
支持自定义LLM,通过代理模式处理自然语言到Pandas代码的转换,内置错误重试机制:
from langchain.agents import create_pandas_dataframe_agent from langchain.llms import HuggingFacePipeline # 加载本地Llama2-Chat模型 llm = HuggingFacePipeline.from_model_id( model_id="meta-llama/Llama-2-13b-chat-hf", task="text-generation", pipeline_kwargs={"temperature": 0.1, "max_new_tokens": 512} ) agent = create_pandas_dataframe_agent(llm, df, verbose=True) agent.run("What is the size of the dataframe")
2. DataChat
轻量工具,专注表格数据的自然语言处理,支持本地模型,可直接输出分析结果或代码,适配多数开源LLM。
3. OpenInterpreter
支持多模态代码生成,能处理DataFrame分析任务,兼容任意本地/云端LLM,自动执行代码并返回结果。
内容的提问来源于stack exchange,提问作者Birender Singh
相关产品推荐
相关产品推荐

