You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy未将所有结果写入CSV的原因排查求助

Hey there, let's break down why you're seeing such a big discrepancy between Scrapy's StatsCollector count and your CSV output, plus address whether errback should catch all those exceptions.

Why the Stats vs CSV Difference?

There are a few common culprits here:

  1. Retries are inflating the Stats count
    By default, Scrapy enables retries (RETRY_ENABLED=True, RETRY_TIMES=2). That means every failed URL gets retried up to 2 additional times—and each retry failure gets counted separately in StatsCollector. But if your errback logic only logs each unique URL once (maybe you're using a set to avoid duplicates, or just writing the first failure), your CSV will only show one entry per URL, even though Stats counts 3 failures for that same URL. This is the most likely reason for a 40k vs 8k gap.

  2. Some exceptions never trigger your errback
    Not all Scrapy exceptions make it to your errback function. Here are a few scenarios where this happens:

  • Middleware-level filtering: If a custom middleware or built-in one (like RobotsTxtMiddleware) blocks a request before it reaches the downloader, the exception might be caught and handled internally without triggering your errback.
  • Unbound errbacks: If you're using CrawlSpider, make sure your Rule objects include the errback parameter. Or if you're generating requests manually in start_requests(), double-check that every Request has errback=self.your_errback_func set—missed requests won't send exceptions to your CSV.
  • Early-stage failures: DNS resolution errors or connection timeouts that happen before the downloader fully processes the request might not propagate to errback if your setup doesn't handle them (though most should, but it's worth checking debug logs).
  1. CSV write conflicts are causing data loss
    If you're writing directly to the CSV file from your errback without handling concurrency, Scrapy's multi-threaded nature can cause race conditions. Multiple errback calls writing to the same file at the same time can overwrite lines or drop entries entirely. Using a thread-safe approach (like a Scrapy Item Pipeline with CsvItemExporter) instead of direct file writes in errback fixes this.

Should errback Catch All Exceptions?

In theory, yes—if configured correctly. To make sure your errback captures as many exceptions as possible:

  • Bind errbacks to every request: Whether you're using start_requests() or CrawlSpider rules, ensure every request has an errback attached.
  • Catch broad exception types: Instead of handling specific exceptions one by one, use a general except Exception as e block in your errback to catch all unexpected errors, then log the exception type and message alongside the URL.
  • Check for middleware interference: Disable custom middlewares temporarily to see if the CSV count increases—if it does, a middleware is swallowing exceptions before they reach errback.

Quick Troubleshooting Steps

  • Set LOG_LEVEL=DEBUG and look for exceptions that aren't showing up in your CSV—this will tell you which errors are slipping through.
  • Temporarily set RETRY_TIMES=0 to disable retries and compare Stats vs CSV counts. If the gap shrinks drastically, retries were the main issue.
  • Switch to using an Item Pipeline for CSV writing instead of direct file writes in errback to eliminate concurrency issues.

内容的提问来源于stack exchange,提问作者Josh Reback

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:07:41