You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Crawl4AI下载Humdata网站PDF文件时遇'int' object has no attribute 'strip'错误的解决求助

Crawl4AI下载Humdata网站PDF文件时遇'int' object has no attribute 'strip'错误的解决求助

大家好,我最近尝试用Crawl4AI工具从Humdata的一个PDF数据集页面批量下载文档,这个网站需要JavaScript渲染才能获取下载链接,我基于官方文档修改了代码,但运行时遇到了一个错误,卡在这里不知道怎么解决,希望能得到大家的帮助。

我修改后的代码如下(主要调整了js_code变量来定位下载按钮):

from crawl4ai.async_configs import BrowserConfig, CrawlerRunConfig
from crawl4ai import AsyncWebCrawler
import os, asyncio
from pathlib import Path

async def download_multiple_files(url: str, download_path: str):
    config = BrowserConfig(accept_downloads=True, downloads_path=download_path)
    async with AsyncWebCrawler(config=config) as crawler:
        run_config = CrawlerRunConfig(
            js_code="""
                const downloadLinks = document.querySelectorAll(
                  "a[title='Download']",
                  document,
                  null,
                  XPathResult.ORDERED_NODE_SNAPSHOT_TYPE,
                  null
                );
                for (const link of downloadLinks) {
                  link.click();
                }
            """,
            wait_for=10  # Wait for all downloads to start
        )
        result = await crawler.arun(url=url, config=run_config)

        if result.downloaded_files:
            print("Downloaded files:")
            for file in result.downloaded_files:
                print(f"- {file}")
        else:
            print("No files downloaded.")

# Usage
download_path = os.path.join(Path.cwd(), "Downloads")
os.makedirs(download_path, exist_ok=True)

asyncio.run(download_multiple_files("https://data.humdata.org/dataset/repository-for-pdf-files", download_path))

运行后抛出了以下错误:

× Unexpected error in _crawl_web at line 1551 in _crawl_web (.venv\Lib\site-packages\crawl4ai\async_crawler_strategy.py):
  Error: Wait condition failed: 'int' object has no attribute 'strip'

  Code context:
  1546                   try:
  1547                       await self.smart_wait(
  1548                           page, config.wait_for, timeout=config.page_timeout
  1549                       )

  1550                   except Exception as e:

  1551 →                     raise RuntimeError(f"Wait condition failed: {str(e)}")

  1552

  1553               # Update image dimensions if needed

  1554               if not self.browser_config.text_mode:

  1555                   update_image_dimensions_js = load_js_script("update_image_dimensions")

  1556                   try:

另外我已经确认过,目标网站的robots.txt允许爬取这个数据集页面,内容如下:

User-agent: *
Disallow: /dataset/rate/
Disallow: /revision/
Disallow: /dataset/*/history
Disallow: /api/
Disallow: /user/*/api-tokens
Crawl-Delay: 10

请问有没有人遇到过类似的问题?或者能帮我分析下这个错误的原因,以及怎么修改代码才能成功下载所有PDF文件呢?

备注:内容来源于stack exchange,提问作者franjefriten

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 11:28:06