You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何利用concurrent.futures实现URL及嵌套页面URL的并发采集?

递归采集网页URL并通过并发缩短耗时

我需要采集网页中锚标签href属性对应的URL,并且对每个获取到的URL重复执行该采集操作,同时希望通过并发机制缩短URL抓取的耗时。以下是实现代码:

import concurrent.futures
from urllib.parse import urlsplit

import requests
from bs4 import BeautifulSoup


def get_href_from_url(url):
    try:
        response = requests.get(url)
        response.raise_for_status()  # 检查HTTP响应是否存在错误
        parts = urlsplit(url)
        base = "{0.netloc}".format(parts)
        strip_base = base.replace("www.", "")
        base_url = "{0.scheme}://{0.netloc}".format(parts)
        # 提取当前URL的路径部分,用于拼接相对路径
        path = url[:url.rfind('/') + 1] if '/' in parts.path else url
        soup = BeautifulSoup(response.text, 'html.parser')
        href_values = set()  # 用集合存储URL,自动去重
        regex = r'.*-c?([0-9]+).html'
        for link in soup.find_all('a'):
            anchor = link.attrs["href"] if "href" in link.attrs else ''
            if anchor.startswith('/'):
                # 处理根路径开头的相对链接
                local_link = base_url + anchor
                href_values.add(local_link)
            elif strip_base in anchor:
                # 处理包含当前域名的链接
                href_values.add(anchor)
            elif not anchor.startswith('http'):
                # 处理相对路径链接
                local_link = path + anchor
                href_values.add(local_link)

        return href_values
    except Exception as e:
        print(f"处理URL {url}时出错: {e}")
        return []


def follow_nested_urls(seed_url):
    visited_urls = set()  # 记录已访问的URL,避免重复抓取
    urls_to_visit = [seed_url]  # 待访问的URL列表

    while len(urls_to_visit):
        current_url = urls_to_visit.pop()
        if current_url in visited_urls:
            continue

        visited_urls.add(current_url)
        # 创建线程池,并发处理请求
        with concurrent.futures.ThreadPoolExecutor(max_workers=20) as executor:
            href_values = get_href_from_url(current_url)

        # 筛选出以http开头的有效URL
        nested_urls = [url for url in href_values if url.startswith('http')]

        urls_to_visit.extend(nested_urls)

        # 输出已访问URL和统计信息
        print(f"已访问URL: {current_url}")
        print(f"已访问总数: {len(visited_urls)}")


if __name__ == "__main__":
    seed_url = "https://www.tradeindia.com/"  # 替换为你需要的起始URL
    follow_nested_urls(seed_url)

代码核心功能说明

  • get_href_from_url:请求目标URL,解析页面所有锚标签的href属性,将各类相对路径转换为绝对路径,返回去重后的有效URL集合;请求或解析出错时打印错误信息并返回空列表。
  • follow_nested_urls:以种子URL为起点,通过维护待访问列表和已访问集合实现递归遍历;使用线程池并发处理URL请求,提升抓取效率,同时输出访问进度统计。

内容的提问来源于stack exchange,提问作者Sam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 08:30:30