LangChain Pandas Agent无法调用python_repl_ast执行Pandas命令问题
问题:LangChain Pandas Agent执行利润均值查询陷入递归循环
问题背景
使用LangChain的Pandas Agent处理超市销售数据集(sales.csv)时,执行"查询利润均值"的指令陷入递归循环。Agent无法正确调用python_repl_ast工具执行Pandas代码,修改该工具描述后问题依旧,且已确认python_repl_ast初始化时已将数据框加载到本地变量。
复现代码
import os os.environ['OPENAI_API_KEY'] = 'sk-xxx' from langchain.agents import create_pandas_dataframe_agent from langchain.chat_models import ChatOpenAI from langchain.llms import OpenAI import pandas as pd df = pd.read_csv('sales.csv') llm = ChatOpenAI(temperature=0.0,model_name='gpt-3.5-turbo') pd_agent = create_pandas_dataframe_agent(llm, df, verbose=True) pd_agent.run("what is the mean of the profit?")
Agent循环日志
> Entering new chain... Thought: We need to calculate the profit first by subtracting the cogs from the total, and then find the mean of the profit. Action: Calculate the profit and find the mean using pandas. Action Input: ---- df['Profit'] = df['Total'] - df['cogs'] df['Profit'].mean() ---- Observation: Calculate the profit and find the mean using pandas. is not a valid tool, try another one. Thought:I need to use python_repl_ast to execute the code. Action: Calculate the profit and find the mean using pandas. Action Input: `python_repl_ast` ---- df['Profit'] = df['Total'] - df['cogs'] df['Profit'].mean() ---- Observation: Calculate the profit and find the mean using pandas. is not a valid tool, try another one. Thought:I need to use `python` instead of `python_repl_ast`. Action: Calculate the profit and find the mean using pandas. Action Input: `python` ---- import pandas as pd df = pd.read_csv('filename.csv') df['Profit'] = df['Total'] - df['cogs'] df['Profit'].mean() ---- . . . Observation: Calculate the profit and find the mean using pandas. is not a valid tool, try another one. Thought: > Finished chain.
最终提示:'Agent stopped due to iteration limit or time limit.'
工具描述修改情况
- 原描述:
A Python shell. Use this to execute python commands. Input should be a valid python command. When using this tool, sometimes output is abbreviated - make sure it does not look abbreviated before using it in your answer. - 修改后描述:
A Python shell. Use this to execute python commands and profit, mean calculation using pandas. Input should be a valid python command. When using this tool, sometimes output is abbreviated - make sure it does not look abbreviated before using it in your answer.
问题原因分析
- Agent输出格式错误:从日志可见,Agent将自然语言描述("Calculate the profit and find the mean using pandas.")作为工具名称传入,而非指定
python_repl_ast作为Action名称,违反了LangChain工具调用的格式要求(正确格式应为Action为工具名,Action Input为对应代码)。 - 工具描述修改无效:修改工具描述未解决LLM输出格式偏差的核心问题,反而可能增加冗余信息干扰模型判断。
- 模型格式理解偏差:gpt-3.5-turbo在零样本场景下,对Pandas Agent的工具调用格式理解不足,尤其当任务需要多步计算时,容易生成不符合要求的输出。
解决建议
- 明确格式引导提示:在查询指令中补充格式示例,强制LLM输出正确的工具调用格式:
pd_agent.run(""" what is the mean of the profit? 请严格按照以下格式输出: Thought: [你的思考内容] Action: python_repl_ast Action Input: [要执行的Python代码] """) - 启用错误处理参数:创建Agent时设置
handle_parsing_errors=True,让Agent自动修正格式错误;同时可调整max_iterations避免过早终止:pd_agent = create_pandas_dataframe_agent( llm, df, verbose=True, handle_parsing_errors=True, max_iterations=10 ) - 预计算简化任务:提前在DataFrame中生成Profit列,减少Agent的计算步骤:
df['Profit'] = df['Total'] - df['cogs'] pd_agent.run("what is the mean of the Profit column?") - 切换至GPT-4模型:GPT-4对工具调用格式的理解精度更高,能有效降低这类格式错误问题。
内容的提问来源于stack exchange,提问作者Jeevan prakash
相关产品推荐
相关产品推荐

