使用Crawl4AI下载Humdata网站PDF文件时遇'int' object has no attribute 'strip'错误的解决求助
Crawl4AI下载Humdata网站PDF文件时遇'int' object has no attribute 'strip'错误的解决求助
大家好,我最近尝试用Crawl4AI工具从Humdata的一个PDF数据集页面批量下载文档,这个网站需要JavaScript渲染才能获取下载链接,我基于官方文档修改了代码,但运行时遇到了一个错误,卡在这里不知道怎么解决,希望能得到大家的帮助。
我修改后的代码如下(主要调整了js_code变量来定位下载按钮):
from crawl4ai.async_configs import BrowserConfig, CrawlerRunConfig from crawl4ai import AsyncWebCrawler import os, asyncio from pathlib import Path async def download_multiple_files(url: str, download_path: str): config = BrowserConfig(accept_downloads=True, downloads_path=download_path) async with AsyncWebCrawler(config=config) as crawler: run_config = CrawlerRunConfig( js_code=""" const downloadLinks = document.querySelectorAll( "a[title='Download']", document, null, XPathResult.ORDERED_NODE_SNAPSHOT_TYPE, null ); for (const link of downloadLinks) { link.click(); } """, wait_for=10 # Wait for all downloads to start ) result = await crawler.arun(url=url, config=run_config) if result.downloaded_files: print("Downloaded files:") for file in result.downloaded_files: print(f"- {file}") else: print("No files downloaded.") # Usage download_path = os.path.join(Path.cwd(), "Downloads") os.makedirs(download_path, exist_ok=True) asyncio.run(download_multiple_files("https://data.humdata.org/dataset/repository-for-pdf-files", download_path))
运行后抛出了以下错误:
× Unexpected error in _crawl_web at line 1551 in _crawl_web (.venv\Lib\site-packages\crawl4ai\async_crawler_strategy.py): Error: Wait condition failed: 'int' object has no attribute 'strip' Code context: 1546 try: 1547 await self.smart_wait( 1548 page, config.wait_for, timeout=config.page_timeout 1549 ) 1550 except Exception as e: 1551 → raise RuntimeError(f"Wait condition failed: {str(e)}") 1552 1553 # Update image dimensions if needed 1554 if not self.browser_config.text_mode: 1555 update_image_dimensions_js = load_js_script("update_image_dimensions") 1556 try:
另外我已经确认过,目标网站的robots.txt允许爬取这个数据集页面,内容如下:
User-agent: * Disallow: /dataset/rate/ Disallow: /revision/ Disallow: /dataset/*/history Disallow: /api/ Disallow: /user/*/api-tokens Crawl-Delay: 10
请问有没有人遇到过类似的问题?或者能帮我分析下这个错误的原因,以及怎么修改代码才能成功下载所有PDF文件呢?
备注:内容来源于stack exchange,提问作者franjefriten
相关产品推荐
相关产品推荐

