You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

以脚本方式运行Scrapy时如何完全禁用其日志?

How to Completely Disable Scrapy Logs in a Script (While Keeping Your Own Logs)

I’ve been in your shoes—trying to run Scrapy as a standalone script and getting swamped with its default logs while wanting to keep your custom logger clean. The key is to target Scrapy’s specific loggers directly without affecting your own setup. Here are the most reliable methods:

Method 1: Disable Scrapy’s Loggers via Python’s logging Module

Scrapy uses Python’s built-in logging system, so we can explicitly disable all loggers that start with "scrapy". Do this before initializing any Scrapy components to catch all early logs:

import logging
from scrapy.crawler import CrawlerProcess
from scrapy.spiders import Spider

# Disable every Scrapy-related logger
for logger_name in logging.root.manager.loggerDict:
    if logger_name.startswith("scrapy"):
        scrapy_logger = logging.getLogger(logger_name)
        # Set level above CRITICAL so no logs are emitted
        scrapy_logger.setLevel(logging.CRITICAL + 1)

# Your custom logger (unchanged)
my_logger = logging.getLogger("my_custom_logger")
my_logger.setLevel(logging.INFO)
handler = logging.StreamHandler()
formatter = logging.Formatter("%(asctime)s - %(name)s - %(levelname)s - %(message)s")
handler.setFormatter(formatter)
my_logger.addHandler(handler)

# Example Spider
class MySpider(Spider):
    name = "example_spider"
    start_urls = ["https://example.com"]

    def parse(self, response):
        my_logger.info("Successfully fetched page content!")
        # Your parsing logic here

# Run the crawler
process = CrawlerProcess()
process.crawl(MySpider)
process.start()

Method 2: Combine with Scrapy’s Settings for Extra Safety

For added insurance, pair the above with Scrapy’s built-in logging-disabling settings. This ensures even edge-case logs (like initial setup messages) are suppressed:

import logging
from scrapy.crawler import CrawlerProcess
from scrapy.spiders import Spider

# Your custom logger setup (same as before)
my_logger = logging.getLogger("my_custom_logger")
my_logger.setLevel(logging.INFO)
handler = logging.StreamHandler()
formatter = logging.Formatter("%(asctime)s - %(name)s - %(levelname)s - %(message)s")
handler.setFormatter(formatter)
my_logger.addHandler(handler)

class MySpider(Spider):
    name = "example_spider"
    start_urls = ["https://example.com"]

    def parse(self, response):
        my_logger.info("Page parsed successfully")

# Configure CrawlerProcess with strict logging rules
process = CrawlerProcess(settings={
    "LOG_ENABLED": False,          # Disable Scrapy's core logging
    "LOG_LEVEL": "CRITICAL",       # Extra safety fallback
    "LOG_STDOUT": False,           # Prevent Scrapy from hijacking stdout
})

process.crawl(MySpider)
process.start()

Why Your Previous Attempts Might Have Failed

  • You probably modified logging settings after initializing Scrapy components (like CrawlerProcess), so early logs were already emitted.
  • Adjusting only the root logger or log level isn’t enough—Scrapy uses dozens of nested loggers (e.g., scrapy.core.engine, scrapy.downloader) that need to be targeted explicitly.

内容的提问来源于stack exchange,提问作者Jimmy Sanchez

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 07:50:30