You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用CrewAI ScrapeWebsiteTool批量爬取多URL遇参数错误求助

CrewAI批量爬取网页问题解决方案

错误原因分析

  1. StructuredTool._run() missing config错误:CrewAI的StructuredTool在执行时会自动传入config参数,若自定义工具的_run方法未声明该参数,就会触发此错误。即使在Task中手动传config,工具本身不接收也没用。
  2. 重复爬取同一网站:大概率是批量处理逻辑中未动态更新URL,或Task的inputs未正确绑定每个URL,导致工具始终使用初始值。

解决方案

方案一:自定义符合规范的批量爬取工具

通过继承StructuredTool,确保_run方法包含config参数,内部循环调用单个ScrapeWebsiteTool完成批量爬取:

from crewai_tools import ScrapeWebsiteTool
from crewai import StructuredTool

class BatchScrapeTool(StructuredTool):
    name: str = "批量网页爬取工具"
    description: str = "一次性爬取多个URL的网页内容,返回每个URL的结果汇总"

    def _run(self, urls: list[str], config=None) -> str:
        """核心执行方法,必须包含config参数"""
        scrape_results = []
        for url in urls:
            # 初始化单个爬取工具
            single_scraper = ScrapeWebsiteTool(url=url)
            # 调用单个工具的_run方法,传入config
            content = single_scraper._run(config=config)
            scrape_results.append(f"=== URL: {url} ===\n{content}\n")
        return "\n".join(scrape_results)

# 工具使用示例
from crewai import Task, Agent, Crew

# 创建Agent
scraper_agent = Agent(
    role="网页数据提取专员",
    goal="准确爬取指定URL的全部可见内容",
    backstory="熟练使用CrewAI工具进行网页数据采集",
    tools=[BatchScrapeTool()]
)

# 创建批量爬取Task
batch_task = Task(
    description="爬取以下URL的内容:{urls}",
    agent=scraper_agent,
    inputs={"urls": ["https://example.com", "https://example.org"]}
)

# 执行任务
crew = Crew(agents=[scraper_agent], tasks=[batch_task])
final_result = crew.kickoff()
print(final_result)

方案二:循环创建独立Task(更简单易维护)

无需自定义工具,直接为每个URL创建单独的Task,从根源避免批量逻辑错误:

from crewai_tools import ScrapeWebsiteTool
from crewai import Task, Agent, Crew

# 初始化单个爬取工具
single_scraper = ScrapeWebsiteTool()

# 创建Agent
scraper_agent = Agent(
    role="网页数据提取专员",
    goal="准确爬取指定URL的全部可见内容",
    backstory="熟练使用CrewAI工具进行网页数据采集",
    tools=[single_scraper]
)

# 待爬取URL列表
target_urls = ["https://example.com", "https://example.org", "https://example.net"]

# 循环生成Task
task_list = []
for index, url in enumerate(target_urls):
    task = Task(
        description=f"爬取URL:{url} 的完整网页内容",
        agent=scraper_agent,
        inputs={"url": url},
        output_file=f"爬取结果_{index}.txt"  # 可选:将每个URL结果保存到单独文件
    )
    task_list.append(task)

# 执行所有Task
crew = Crew(agents=[scraper_agent], tasks=task_list)
results = crew.kickoff()

# 输出所有结果
for res in results:
    print(res)

关键注意点

  • 自定义工具时,_run方法必须保留config参数,即使代码中未使用,否则会触发参数缺失错误。
  • 批量爬取时,确保每个URL都被正确传入Task的inputs,避免因变量引用错误导致重复爬取同一站点。

内容的提问来源于stack exchange,提问作者Gs can't

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 23:06:10