You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用PHP验证9万条URL的HTTP状态码,寻求提速方案

Laravel批量URL状态码验证优化方案

我手里有9万条URL,需要验证每条的HTTP状态码并更新到数据库对应条目中。目前基于Laravel实现的代码运行速度极慢,测试下来要6-7小时,想通过并发请求来提速。尝试用Laravel Http客户端实现并发时遇到了困难,附上现有实现代码和尝试的并发代码片段,恳请给出优化建议。

原有实现代码

public function handle()
{
    
    DB::table('internal_links')->whereNull('status')->orderBy('id')->chunk(5, function($links) {
       
       foreach($links as $link) {
            stream_context_get_default([
                'http' => ['method' => 'HEAD']
            ]);
            $status_code = @get_headers($link->href)[0];
            if(!$status_code) {
                $status_code = 404;
            } else {
                $status_code = substr($status_code, 9,3);
            }
            DB::table('internal_links')->where('href', $link->href)->update(['status' => $status_code]);
       }
    });
    return Command::SUCCESS;
}

尝试的并发代码片段

public function handle()
{
    DB::table('internal_links')->distinct('href')->orderBy('id')->chunk(50, function($urls) {
        // 怎么把循环$urls里的每个$url放到$pool->head($url)里?
        $responses = Http::pool(fn (Pool $pool) => [
            $pool->head('http://url1'),
            $pool->head('http://url2')
        ]);
    });
    return Command::SUCCESS;
}

优化方案

1. 正确实现Laravel Http客户端并发请求

核心是批量构建请求任务,同时保留URL与原始数据的关联,拿到响应后批量更新数据库,避免单次请求+单次更新的低效模式。

修改后的完整代码:

use Illuminate\Http\Client\Pool;
use Illuminate\Support\Facades\Http;

public function handle()
{
    // 调整chunk大小,建议从100开始测试,根据服务器性能和目标网站限流策略调整
    DB::table('internal_links')->whereNull('status')->orderBy('id')->chunk(100, function($links) {
        // 构建带关联信息的请求数组
        $requestMap = [];
        foreach ($links as $link) {
            // 保存href作为键,方便后续匹配响应
            $requestMap[$link->href] = fn(Pool $pool) => $pool->head($link->href)->timeout(10);
        }

        // 发起并发请求
        $responses = Http::pool(fn(Pool $pool) => array_map(function($request) use ($pool) {
            return $request($pool);
        }, $requestMap));

        // 整理批量更新数据
        $updateData = [];
        foreach ($responses as $href => $response) {
            // 处理各种异常情况,返回对应状态码
            if ($response->failed()) {
                $statusCode = $response->status() ?: 404;
            } else {
                $statusCode = $response->status();
            }
            $updateData[] = [
                'href' => $href,
                'status' => $statusCode
            ];
        }

        // 使用upsert批量更新,比循环单个update效率提升数倍
        if (!empty($updateData)) {
            DB::table('internal_links')->upsert(
                $updateData,
                ['href'], // 唯一标识字段,需确保href是唯一的或有唯一索引
                ['status']
            );
        }
    });

    return Command::SUCCESS;
}

2. 关键优化点说明

  • 并发请求批量处理:用Http::pool一次性发起多个HEAD请求,替代串行请求,大幅减少等待时间。
  • 批量数据库更新:用upsert替代循环update,减少数据库连接次数和查询开销。
  • 超时控制:给每个请求设置超时时间(如10秒),避免单个请求卡住拖慢整个批次。
  • 错误处理:去掉@错误抑制,通过$response->failed()精准判断请求状态,返回对应状态码。

3. 额外性能优化建议

  • 数据库索引优化:
    • 给internal_links表的href字段添加唯一索引,无论是查询还是更新都会大幅提速。
    • 给status字段添加普通索引,加快whereNull('status')的筛选速度。
  • 调整并发批次大小:
    • 不要盲目增大chunk size,建议从100开始测试,逐步调整到服务器和目标网站能承受的最大值,避免触发限流或耗尽服务器资源。
  • 队列异步处理(可选):
    • 如果并发请求压力仍大,可以将每个chunk的任务封装成Laravel队列任务,启动多个worker同时处理不同批次,进一步提升整体效率。
  • 失败重试机制:
    • 对请求失败的URL记录到单独表中,后续通过命令重试,避免遗漏。

内容的提问来源于stack exchange,提问作者TheBigK

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 19:21:06