You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何向concurrent.futures批量传入链接而非逐个传入?

批量提交多链接爬取任务到ThreadPoolExecutor

你的核心需求是将所有初始链接的店铺任务一次性批量提交给线程池,而非逐个处理每个初始链接。当前代码是循环每个初始链接,单独提交该链接下的店铺任务,我们可以通过合并所有店铺任务的生成器,实现批量提交。

修改后的完整代码

import requests
from bs4 import BeautifulSoup
import concurrent.futures
# 请补充你的请求头定义
headers = {"User-Agent": "Mozilla/5.0 ..."}

links = [
    'first link',
    'second link',
    'third link',
]

def get_links(link):
    while True:
        res = requests.get(link, headers=headers)
        soup = BeautifulSoup(res.text, "html.parser")
        for item in soup.select("[data-testid='serp-ia-card'] [class*='businessName'] a[href^='/biz/'][name]"):
            shop_name = item.get_text(strip=True)
            shop_link = item.get('href')
            # 补全相对链接为绝对链接,避免请求失败
            if not shop_link.startswith('http'):
                shop_link = f"https://example.com{shop_link}"
            yield shop_name, shop_link
        
        next_page = soup.select_one("a.next-link[aria-label='Next']")
        if not next_page:
            return
        link = next_page.get("href")
        # 补全下一页链接
        if not link.startswith('http'):
            link = f"https://example.com{link}"

def get_content(shop_name, shop_link):
    try:
        res = requests.get(shop_link, headers=headers)
        res.raise_for_status()  # 主动抛出HTTP请求错误
        soup = BeautifulSoup(res.text, "html.parser")
        phone = soup.select_one("p:-soup-contains('Phone number') + p").get_text(strip=True)
    except (AttributeError, TypeError, requests.exceptions.RequestException):
        phone = ""
    return shop_name, shop_link, phone

if __name__ == '__main__':
    with concurrent.futures.ThreadPoolExecutor(max_workers=6) as executor:
        # 合并所有初始链接的店铺任务生成器,避免一次性加载大量数据到内存
        all_shop_tasks = (elem for link in links for elem in get_links(link))
        
        # 批量提交所有任务到线程池
        future_to_url = {executor.submit(get_content, *elem): elem for elem in all_shop_tasks}
        
        # 按任务完成顺序处理结果
        for future in concurrent.futures.as_completed(future_to_url):
            try:
                shop_name, shop_link, phone = future.result()
                print(shop_name, shop_link, phone)
            except Exception as e:
                # 捕获单个任务的执行异常,避免程序崩溃
                elem = future_to_url[future]
                print(f"处理任务 {elem} 时出错: {e}")

关键修改说明

  1. 合并任务生成器
    使用生成器表达式(elem for link in links for elem in get_links(link)),遍历所有初始链接并合并每个链接下的店铺任务,无需一次性将所有店铺数据加载到内存,适合处理大量数据的场景。

  2. 批量提交任务
    直接基于合并后的迭代器创建future_to_url字典,一次性将所有店铺任务提交给线程池,让线程池更高效地调度任务,避免逐个处理初始链接时的资源闲置。

  3. 异常增强处理

    • 补全相对链接为绝对链接,解决请求无效路径的问题
    • 在get_content中新增网络请求异常捕获,覆盖HTTP错误、连接失败等场景
    • 结果处理阶段增加异常捕获,确保单个任务失败不会导致整个程序终止

内容的提问来源于stack exchange,提问作者robots.txt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 19:01:18