如何在Laravel中用Spatie SitemapGenerator排除非200状态码URL
在Laravel中用Spatie SitemapGenerator排除非200状态码URL的优化方案
问题场景
用Laravel+Spatie SitemapGenerator生成站点地图,想过滤掉状态码不是200的URL,已经在guzzle_options里设了ALLOW_REDIRECTS => false,但302跳转的URL还是被收录。试过在hasCrawled回调里调用$url->getStatusCode(),但根本没用。后来临时用了个方案,虽然能生效,但3000+URL的生成时间从5分钟涨到17分钟,性能拉胯,急需更好的办法。
现有配置与代码
config/sitemap.php
return [ 'guzzle_options' => [ RequestOptions::COOKIES => true, RequestOptions::CONNECT_TIMEOUT => 0, RequestOptions::TIMEOUT => 0, RequestOptions::ALLOW_REDIRECTS => false, ], 'execute_javascript' => false, 'chrome_binary_path' => null, 'crawl_profile' => Profile::class, ];
SitemapController.php
use Spatie\Sitemap\SitemapGenerator; use Spatie\Sitemap\Tags\Url; SitemapGenerator::create(config('app.url')) ->hasCrawled(function (Url $url) { // 排除带查询字符串的URL if (strpos($url->url, '?') !== false) { return; } // 这里怎么检查状态码? return $url; }) ->writeToFile(public_path('sitemap.xml'));
临时方案(性能极差)
下面这个方案能过滤非200的URL,但因为每个URL都额外发起一次HEAD请求,直接导致性能暴跌:
$client = new \GuzzleHttp\Client(); $crawled_urls = []; SitemapGenerator::create(config('app.url')) ->shouldCrawl(function (UriInterface $url) use($client, &$crawled_urls) { if($url->getQuery() != ''){ return false; } $endpoint = (string)$url; if(isset($crawled_urls[$endpoint])){ return $crawled_urls[$endpoint]; } $response = $client->head($endpoint, ['allow_redirects' => false]); $crawled_urls[$endpoint] = ($response->getStatusCode() == 200); return $crawled_urls[$endpoint]; })->writeToFile(public_path('sitemap.xml'));
优化方案
Spatie的爬虫在爬取过程中已经拿到了每个URL的响应状态码,没必要额外发请求。直接用官方支持的自定义Profile或者CrawlObserver就能解决,性能和原生爬虫一致。
方法1:自定义Crawl Profile(推荐)
直接在爬虫阶段就过滤掉非200的URL,全程只爬一次,没有额外开销:
- 新建自定义Profile类:
// app/Services/CustomCrawlProfile.php namespace App\Services; use Spatie\Crawler\CrawlProfile; use Psr\Http\Message\ResponseInterface; use Psr\Http\Message\UriInterface; class CustomCrawlProfile extends CrawlProfile { public function shouldCrawl(UriInterface $url): bool { // 先排除带查询字符串的URL if ($url->getQuery() !== '') { return false; } return true; } public function shouldCrawlResponse(ResponseInterface $response): bool { // 只保留状态码为200的响应 return $response->getStatusCode() === 200; } }
- 修改配置文件config/sitemap.php,替换默认的Profile:
'crawl_profile' => App\Services\CustomCrawlProfile::class,
- 简化控制器代码:
use Spatie\Sitemap\SitemapGenerator; SitemapGenerator::create(config('app.url')) ->writeToFile(public_path('sitemap.xml'));
方法2:自定义CrawlObserver(进阶)
如果需要更灵活的处理逻辑,可以直接监听爬虫的响应,收集符合条件的URL再生成站点地图:
use Spatie\Sitemap\SitemapGenerator; use Psr\Http\Message\UriInterface; use Psr\Http\Message\ResponseInterface; use Spatie\Crawler\CrawlObservers\CrawlObserver; SitemapGenerator::create(config('app.url')) ->setCrawlObserver(new class extends CrawlObserver { protected $validUrls = []; public function crawled(UriInterface $url, ResponseInterface $response, ?UriInterface $foundOnUrl = null) { // 只收集状态码200且无查询字符串的URL if ($response->getStatusCode() === 200 && $url->getQuery() === '') { $this->validUrls[] = (string)$url; } } public function finishedCrawling() { // 生成并写入站点地图 $sitemap = \Spatie\Sitemap\Sitemap::create(); foreach ($this->validUrls as $url) { $sitemap->add(\Spatie\Sitemap\Tags\Url::create($url)); } $sitemap->writeToFile(public_path('sitemap.xml')); } }) ->startCrawling();
方案说明
- 方法1是官方推荐的标准用法,完全利用爬虫原生逻辑,性能最优,代码也最简洁。
- 方法2适合需要对爬取结果做更多自定义处理的场景,同样不会额外发起HTTP请求,性能和原生爬虫一致。
内容的提问来源于stack exchange,提问作者Wilson
相关产品推荐
相关产品推荐

