You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Crawlee报错:找不到request_queues文件,如何解决?

解决PlaywrightCrawler请求队列"文件或目录不存在"错误

问题描述

运行一个需抓取约64000个页面的爬虫时,反复出现ENOENT: no such file or directory错误。即使仅处理1000个页面、设置waitForAllRequestsToBeAdded: true或调整爬虫选项,问题依然存在。

错误信息

ERROR PlaywrightCrawler:AutoscaledPool: runTaskFunction failed.
Error: ENOENT: no such file or directory, open '/project/storage/request_queues/default/ehieKpeBY6Mf39n.json'

爬虫代码

const opts = {
    navigationTimeoutSecs: 3,
    requestHandlerTimeoutSecs: 3,
    maxRequestRetries: 6,
    maxConcurrency: 20
};
const config = new Configuration({
    memoryMbytes: 8000
});
const crawler = new PlaywrightCrawler(opts, config);
crawler.router.addDefaultHandler(handlePage);
const requests = data.map(
    (d) =>
    new Request({
        url: d.url,
        userData: d
    })
);
await crawler.run(requests);

解决方案

1. 修复存储目录的存在性与权限

错误指向的/project/storage/request_queues/default目录可能未创建,或爬虫进程无读写权限:

  • 手动创建完整目录链:mkdir -p /project/storage/request_queues/default
  • 赋予读写权限:chmod -R 755 /project/storage

2. 切换为内存型请求队列

如果不需要持久化请求状态,直接用内存存储替代文件系统,彻底规避文件IO问题:

// 先导入MemoryStorage
import { MemoryStorage } from 'crawlee';

const opts = {
    navigationTimeoutSecs: 3,
    requestHandlerTimeoutSecs: 3,
    maxRequestRetries: 6,
    maxConcurrency: 20,
    // 指定内存存储
    requestQueueOptions: {
        storage: new MemoryStorage()
    }
};

3. 显式指定合法的存储路径

若需要持久化,在Configuration中明确设置存在且有权限的存储目录:

const config = new Configuration({
    memoryMbytes: 8000,
    storageDir: '/your/valid/existing/storage/path'
});

4. 临时降低并发与重试次数

过高的并发或重试可能引发文件系统竞争,临时调低参数排查问题:

  • 将maxConcurrency设为5,maxRequestRetries设为2,测试是否能稳定运行,逐步定位是否为IO竞争导致的文件缺失。

内容的提问来源于stack exchange,提问作者sbrass

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 21:12:42