Scrapy未将所有结果写入CSV的原因排查求助
Hey there, let's break down why you're seeing such a big discrepancy between Scrapy's StatsCollector count and your CSV output, plus address whether errback should catch all those exceptions.
Why the Stats vs CSV Difference?
There are a few common culprits here:
Retries are inflating the Stats count
By default, Scrapy enables retries (RETRY_ENABLED=True,RETRY_TIMES=2). That means every failed URL gets retried up to 2 additional times—and each retry failure gets counted separately in StatsCollector. But if yourerrbacklogic only logs each unique URL once (maybe you're using a set to avoid duplicates, or just writing the first failure), your CSV will only show one entry per URL, even though Stats counts 3 failures for that same URL. This is the most likely reason for a 40k vs 8k gap.Some exceptions never trigger your errback
Not all Scrapy exceptions make it to yourerrbackfunction. Here are a few scenarios where this happens:
- Middleware-level filtering: If a custom middleware or built-in one (like
RobotsTxtMiddleware) blocks a request before it reaches the downloader, the exception might be caught and handled internally without triggering yourerrback. - Unbound errbacks: If you're using
CrawlSpider, make sure yourRuleobjects include theerrbackparameter. Or if you're generating requests manually instart_requests(), double-check that everyRequesthaserrback=self.your_errback_funcset—missed requests won't send exceptions to your CSV. - Early-stage failures: DNS resolution errors or connection timeouts that happen before the downloader fully processes the request might not propagate to
errbackif your setup doesn't handle them (though most should, but it's worth checking debug logs).
- CSV write conflicts are causing data loss
If you're writing directly to the CSV file from yourerrbackwithout handling concurrency, Scrapy's multi-threaded nature can cause race conditions. Multipleerrbackcalls writing to the same file at the same time can overwrite lines or drop entries entirely. Using a thread-safe approach (like a Scrapy Item Pipeline withCsvItemExporter) instead of direct file writes inerrbackfixes this.
Should errback Catch All Exceptions?
In theory, yes—if configured correctly. To make sure your errback captures as many exceptions as possible:
- Bind errbacks to every request: Whether you're using
start_requests()orCrawlSpiderrules, ensure every request has anerrbackattached. - Catch broad exception types: Instead of handling specific exceptions one by one, use a general
except Exception as eblock in yourerrbackto catch all unexpected errors, then log the exception type and message alongside the URL. - Check for middleware interference: Disable custom middlewares temporarily to see if the CSV count increases—if it does, a middleware is swallowing exceptions before they reach
errback.
Quick Troubleshooting Steps
- Set
LOG_LEVEL=DEBUGand look for exceptions that aren't showing up in your CSV—this will tell you which errors are slipping through. - Temporarily set
RETRY_TIMES=0to disable retries and compare Stats vs CSV counts. If the gap shrinks drastically, retries were the main issue. - Switch to using an Item Pipeline for CSV writing instead of direct file writes in
errbackto eliminate concurrency issues.
内容的提问来源于stack exchange,提问作者Josh Reback

