You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中并发下载CommonCrawl的多个Warc大文件?

批量从CommonCrawl下载WARC文件的脚本优化

我看到你正在尝试从CommonCrawl数据库批量获取WARC文件,脚本框架已经搭好了,但后半部分还没完成。我来帮你补全并优化这段代码,让它能顺利完成索引查询和文件下载任务:

首先先贴出你的原始脚本片段:

import json
import urllib.request
from urllib.error import HTTPError
from src.Util import rooted

with open(rooted('data/alexa.txt'), 'r') as alexa:
    for i, url in enumerate(alexa):
        if i % 1000 == 0:
            try:
                request = 'http://index.commoncrawl.org/CC-MAIN-2018-13-index?url={search}*&output=json' \
                    .format(search=url.rstrip())
                page = urllib.request.urlopen(request)
                for line in ...

下面是补全并优化后的完整脚本,我加了详细注释说明每个步骤:

import json
import urllib.request
from urllib.error import HTTPError, URLError
from src.Util import rooted
import os

# 提前创建WARC文件保存目录,避免路径不存在报错
save_dir = rooted('data/warc_files')
os.makedirs(save_dir, exist_ok=True)

with open(rooted('data/alexa.txt'), 'r') as alexa:
    for i, url in enumerate(alexa):
        url = url.rstrip()  # 清除URL首尾的换行符和空白字符
        if not url:
            continue  # 跳过空行,避免无效查询
        
        # 保留你原来的每1000条URL查询一次索引的逻辑
        if i % 1000 == 0:
            try:
                # 修正请求URL的转义字符:代码里不需要用&,直接用&即可
                request_url = f'http://index.commoncrawl.org/CC-MAIN-2018-13-index?url={url}*&output=json'
                print(f"正在查询索引: {request_url}")
                
                # 给请求加超时限制,避免长时间卡住
                with urllib.request.urlopen(request_url, timeout=10) as page:
                    # CommonCrawl的索引输出是每行一个JSON对象,逐行处理
                    for line in page:
                        try:
                            # 解析单行JSON数据
                            record = json.loads(line.decode('utf-8'))
                            # 提取WARC文件的下载链接
                            warc_url = record.get('filename')
                            if not warc_url:
                                print(f"未找到{url}对应的WARC文件链接")
                                continue
                            
                            # 生成本地保存路径,用WARC文件名作为本地文件名
                            warc_filename = os.path.basename(warc_url)
                            save_path = os.path.join(save_dir, warc_filename)
                            
                            # 如果文件已经存在,跳过下载节省时间
                            if os.path.exists(save_path):
                                print(f"文件已存在,跳过: {save_path}")
                                continue
                            
                            # 开始下载WARC文件
                            print(f"开始下载: {warc_url} -> {save_path}")
                            urllib.request.urlretrieve(warc_url, save_path)
                            print(f"下载完成: {save_path}")
                            
                        except json.JSONDecodeError:
                            print(f"JSON解析失败,行内容: {line}")
                        except HTTPError as e:
                            print(f"下载WARC文件出错: {e.code} {e.reason}")
                        except URLError as e:
                            print(f"网络连接错误: {e.reason}")
            
            except HTTPError as e:
                print(f"查询索引出错: {e.code} {e.reason},请求URL: {request_url}")
            except URLError as e:
                print(f"网络连接失败: {e.reason},请求URL: {request_url}")
            except Exception as e:
                print(f"未知错误: {str(e)},请求URL: {request_url}")

几个关键优化点说明:

  • 修正了请求URL中的&为&,代码里不需要使用HTML转义字符
  • 新增目录创建逻辑,确保WARC文件有合法的保存路径
  • 增强了异常处理,覆盖JSON解析、HTTP错误、网络错误等多种场景
  • 添加文件存在检查,避免重复下载已有的WARC文件
  • 加入请求超时设置,防止脚本因网络问题长时间卡住
  • 增加运行日志打印,方便跟踪脚本执行状态

如果需要筛选特定状态码的记录,或者只下载某个时间段的WARC文件,可以根据CommonCrawl索引返回的status、timestamp等字段进一步调整逻辑。

内容的提问来源于stack exchange,提问作者kabeersvohra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 04:15:04