You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Rust tokio实现未知总页数的分页网页并发爬取

实现方案

你当前使用的try_join_all需要预先收集所有待执行的Future,天然不支持动态终止的场景,且你不需要用到futures::executor::block_on,该API仅用于在同步上下文调用异步逻辑,你已经处于tokio的异步运行时上下文中,直接使用流处理即可实现需求。

依赖配置

首先在Cargo.toml中添加futures依赖用于异步流处理:

[dependencies]
tokio = { version = "1.0", features = ["full"] }
futures = "0.3"

示例代码

use futures::stream::{self, StreamExt};

async fn fetch(url: String) -> Result<String, ()> {
    tokio::time::sleep(std::time::Duration::from_millis(100)).await;
    // 模拟页码>=5时返回无内容标记
    if url.contains("page=5") {
        return Ok(String::from("not found"));
    }
    Ok(format!("爬取结果:{}", url))
}

#[tokio::main]
async fn main() {
    let search_url = "https://example.com/?page={page_num}";

    let all_valid_results = stream::unfold(1, |current_page| async move {
        // 动态生成下一页的爬取任务,无需预先指定总页数
        let target_url = search_url.replace("{page_num}", &current_page.to_string());
        let fetch_result = fetch(target_url).await;
        Some((fetch_result, current_page + 1))
    })
    // 限制并发数为3,避免请求频率过高被目标站点封禁
    .buffered(3)
    // 遇到终止条件立刻停止爬取,后续请求不会再发起
    .take_while(|result| async {
        match result {
            Ok(content) => !content.contains("not found"),
            // 若遇到请求错误也可以选择终止,或者调整为返回true继续爬取下一页
            Err(_) => false
        }
    })
    // 收集所有有效爬取结果
    .collect::<Vec<_>>()
    .await;

    // 处理爬取结果
    for res in all_valid_results {
        println!("{}", res.unwrap());
    }
}

逻辑说明

  • stream::unfold用于生成动态的异步任务流,从页码1开始,每执行完当前页的爬取就自动生成下一页的爬取任务,不需要提前知晓总页数
  • 你可以根据需求调整并发控制逻辑:buffered(N)会保持页码返回顺序,buffer_unordered(N)不保持顺序,执行效率更高
  • take_while中可以自定义任意终止条件,除了返回内容匹配之外,你也可以加入状态码判断、返回内容为空判断等规则,满足条件就会立刻终止整个爬取流

内容的提问来源于stack exchange,提问作者orhun

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 00:57:03