You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用CrewAI+Ollama Phi3生成Python代码时触发<|endoftext|>报错

问题:CrewAI搭配Ollama Phi3传递代码给HumanProxy时触发特殊令牌错误

现象

  • 搭配Ollama Phi3使用CrewAI时,无需生成Python代码的场景运行正常
  • 当Developer Agent将包含Python代码的响应传递给HumanProxy Agent时,任务终止并持续抛出令牌编码相关报错

报错信息

2024-05-02 14:24:07,361 - 129651365437888 - manager.py-manager:281 - WARNING: Error in TokenCalcHandler.on_llm_start callback: ValueError("Encountered text corresponding to disallowed special token '<|endoftext|>'.
If you want this text to be encoded as a special token, pass it to allowed_special, e.g. allowed_special={'<|endoftext|>', ...}.
If you want this text to be encoded as normal text, disable the check for this token by passing disallowed_special=(enc.special_tokens_set - {'<|endoftext|>'}).
To disable this check for all special tokens, pass disallowed_special=().
")

运行输出片段

>>> result = crew.kickoff()
 [DEBUG]: == Working Agent: Senior Manager
 [INFO]: == Starting Task: Create python script that scrape the websites and saves the output into the text file fo further processing.
  Manage the work with your crew to answer the request. 
  Check with Human if the response was helpful and if it is not satisfactory then try again. 


> Entering new CrewAgentExecutor chain...
Thought: I need to delegate tasks efficiently among my crew members to create a Python script for web scraping that saves output into a text file.
Action: Delegate work to co-worker
Action Input: {"coworker": "Python Developer", "task": "Create a Python script for web scraping and saving the output into a text file.", "context": "The task involves using libraries like BeautifulS Authorization (BeautifulSoup) or Scrapy to scrape websites. The data extracted should be saved in a structured format, such as JSON or CSV, which can then be processed further."}

> Entering new CrewAgentExecutor chain...
Thought: To complete this task effectively, I need to select a suitable library for web scraping, define the target website URL, specify the data elements to be extracted, and then write Python code that will scrape the desired information and save it in JSON format.

Final Answer:
```python
import requests
from bs4 import BeautifulSoup
import json

# Define the URL of the website you want to scrape
url = 'https://example.com'  # Replace with actual target website URL

def web_scrape(target_url):
    try:
        response = requests.get(target_url)
        soup = BeautifulSoup(response.text, 'html.parser')
        
        # Assuming we want to extract data from elements with a specific class or id
        data_elements = []  # Replace this with the actual logic for selecting data elements
        for element in soup.select('your-selector'):  # Replace 'your-selector' with appropriate CSS selector(s)
            extracted_data = {
                'element_id': element['id'],  # Assuming each element has an id attribute
                'text': element.get_text()
            }
            data_elements.append(extracted_data)
        
        # Save the scraped data into a JSON file
        with open('output.json', 'w') as json_file:
            json.dump(data_elements, json_file, indent=4)
            
    except requests.RequestException as e:
        print(f"An error occurred while making the request: {e}")

# Execute the web scraping function
web_scrape(url)

Replace 'https://example.com' with your actual target website URL and adjust the data extraction logic according to the specific elements you need to scrape from that site.<|end|><|endoftext|>

Finished chain.

import requests
from bs4 import BeautifulSoup
import json

# Define the URL of the website you want to scrape
url = 'https://example.com'  # Replace with actual target website URL

def web_scrape(target_url):
    try:
        response = requests.get(target_url)
        soup = BeautifulSoup(response.text, 'html.parser')
        
        # Assuming we want to extract data from elements with a specific class or id
        data_elements = []  # Replace this with the actual logic for selecting data elements
        for element in soup.select('your-selector'):  # Replace 'your-selector' with appropriate CSS selector(s)
            extracted_data = {
                'element_id': element['id'],  # Assuming each element has an id attribute
                'text': element.get_text()
            }
            data_elements.append(extracted_data)
        
        # Save the scraped data into a JSON file
        with open('output.json', 'w') as json_file:
            json.dump(data_elements, json_file, indent=4)
            
    except requests.RequestException as e:
        print(f"An error occurred while making the request: {e}")

# Execute the web scraping function
web_scrape(url)

Replace 'https://example.com' with your actual target website URL and adjust the data extraction logic according to the specific elements you need to scrape from that site.<|end|><|endoftext|>

## 解决方案
### 核心原因
Phi3模型输出的响应末尾自带`<|end|><|endoftext|>`特殊令牌,HumanProxy Agent的令牌计算模块(TokenCalcHandler)将其判定为非法特殊令牌,触发编码错误。

### 修复方案
#### 方案1:修改LLM令牌编码配置
初始化Ollama模型时,允许`<|endoftext|>`作为合法特殊令牌,或直接禁用特殊令牌检查:
```python
from langchain_community.llms import Ollama

llm = Ollama(
    model="phi3",
    model_kwargs={
        "allowed_special": {"<|endoftext|>"},
        "disallowed_special": ()  # 若仅需放行<|endoftext|>,可改为:enc.special_tokens_set - {'<|endoftext|>'}
    }
)

方案2:清理响应中的特殊令牌

为Developer Agent添加自定义输出解析器,自动移除响应末尾的特殊令牌:

from langchain.schema import BaseOutputParser

class CleanSpecialTokensParser(BaseOutputParser):
    def parse(self, text: str) -> str:
        cleaned_text = text.replace("<|end|>", "").replace("<|endoftext|>", "").strip()
        return cleaned_text

# 为Developer Agent配置该解析器
developer_agent = Agent(
    role="Python Developer",
    llm=llm,
    output_parser=CleanSpecialTokensParser(),
    # 其他Agent配置参数...
)

方案3:禁用TokenCalcHandler

如果不需要令牌统计功能,直接移除Crew的默认回调处理器:

from crewai import Crew, Process

crew = Crew(
    agents=[senior_manager, developer_agent, human_proxy],
    tasks=[task],
    process=Process.sequential,
    callbacks=[]  # 清空回调列表,移除TokenCalcHandler
)

内容的提问来源于stack exchange,提问作者m1k3y3

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 17:14:53