如何在Python中并发下载CommonCrawl的多个Warc大文件?
批量从CommonCrawl下载WARC文件的脚本优化
我看到你正在尝试从CommonCrawl数据库批量获取WARC文件,脚本框架已经搭好了,但后半部分还没完成。我来帮你补全并优化这段代码,让它能顺利完成索引查询和文件下载任务:
首先先贴出你的原始脚本片段:
import json import urllib.request from urllib.error import HTTPError from src.Util import rooted with open(rooted('data/alexa.txt'), 'r') as alexa: for i, url in enumerate(alexa): if i % 1000 == 0: try: request = 'http://index.commoncrawl.org/CC-MAIN-2018-13-index?url={search}*&output=json' \ .format(search=url.rstrip()) page = urllib.request.urlopen(request) for line in ...
下面是补全并优化后的完整脚本,我加了详细注释说明每个步骤:
import json import urllib.request from urllib.error import HTTPError, URLError from src.Util import rooted import os # 提前创建WARC文件保存目录,避免路径不存在报错 save_dir = rooted('data/warc_files') os.makedirs(save_dir, exist_ok=True) with open(rooted('data/alexa.txt'), 'r') as alexa: for i, url in enumerate(alexa): url = url.rstrip() # 清除URL首尾的换行符和空白字符 if not url: continue # 跳过空行,避免无效查询 # 保留你原来的每1000条URL查询一次索引的逻辑 if i % 1000 == 0: try: # 修正请求URL的转义字符:代码里不需要用&,直接用&即可 request_url = f'http://index.commoncrawl.org/CC-MAIN-2018-13-index?url={url}*&output=json' print(f"正在查询索引: {request_url}") # 给请求加超时限制,避免长时间卡住 with urllib.request.urlopen(request_url, timeout=10) as page: # CommonCrawl的索引输出是每行一个JSON对象,逐行处理 for line in page: try: # 解析单行JSON数据 record = json.loads(line.decode('utf-8')) # 提取WARC文件的下载链接 warc_url = record.get('filename') if not warc_url: print(f"未找到{url}对应的WARC文件链接") continue # 生成本地保存路径,用WARC文件名作为本地文件名 warc_filename = os.path.basename(warc_url) save_path = os.path.join(save_dir, warc_filename) # 如果文件已经存在,跳过下载节省时间 if os.path.exists(save_path): print(f"文件已存在,跳过: {save_path}") continue # 开始下载WARC文件 print(f"开始下载: {warc_url} -> {save_path}") urllib.request.urlretrieve(warc_url, save_path) print(f"下载完成: {save_path}") except json.JSONDecodeError: print(f"JSON解析失败,行内容: {line}") except HTTPError as e: print(f"下载WARC文件出错: {e.code} {e.reason}") except URLError as e: print(f"网络连接错误: {e.reason}") except HTTPError as e: print(f"查询索引出错: {e.code} {e.reason},请求URL: {request_url}") except URLError as e: print(f"网络连接失败: {e.reason},请求URL: {request_url}") except Exception as e: print(f"未知错误: {str(e)},请求URL: {request_url}")
几个关键优化点说明:
- 修正了请求URL中的
&为&,代码里不需要使用HTML转义字符 - 新增目录创建逻辑,确保WARC文件有合法的保存路径
- 增强了异常处理,覆盖JSON解析、HTTP错误、网络错误等多种场景
- 添加文件存在检查,避免重复下载已有的WARC文件
- 加入请求超时设置,防止脚本因网络问题长时间卡住
- 增加运行日志打印,方便跟踪脚本执行状态
如果需要筛选特定状态码的记录,或者只下载某个时间段的WARC文件,可以根据CommonCrawl索引返回的status、timestamp等字段进一步调整逻辑。
内容的提问来源于stack exchange,提问作者kabeersvohra
相关产品推荐
相关产品推荐

