You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

重装系统后Scrapy爬虫无输出,日志显示Enabled Item Pipeline: []求助

问题分析

从日志能看出两个核心问题:

  1. 你的目标URL(https://www.quicktransportsolutions.com/carrier/usa-trucking-companies.php)根本没被Scrapy调度,反而请求了网站首页,说明蜘蛛运行逻辑或网站限制导致目标页面未被访问。
  2. 日志里的Enabled item pipelines: []不是故障原因——没有启用管道时,Scrapy会直接把爬取到的内容输出到控制台,不影响数据提取。

修复方案

1. 确保蜘蛛正确运行

你的蜘蛛名称是truckspider,但日志显示bot名称为truckscraper2,需确认运行命令是否正确:

# 必须指定你的蜘蛛名称
scrapy crawl truckspider

如果用runspider命令,要确保指定了正确的脚本文件:

scrapy runspider your_script_filename.py

2. 解除Robots协议限制(若合规)

访问目标网站的robots.txt(https://www.quicktransportsolutions.com/robots.txt),如果发现它禁止爬取/carrier/路径,修改项目的settings.py:

# 关闭Robots协议遵守,注意需确保爬取行为合规
ROBOTSTXT_OBEY = False

3. 优化CSS选择器(适配页面结构)

原脚本的选择器[class="col-md-4 column"]可能因页面空格或结构变化无法匹配元素,替换为更健壮的选择器,同时处理空值:

import scrapy

class TruckspiderSpider(scrapy.Spider):
    name = 'truckspider'
    allowed_domains = ['www.quicktransportsolutions.com']
    start_urls = ['https://www.quicktransportsolutions.com/carrier/usa-trucking-companies.php']

    def parse(self, response):
        # 用类选择器匹配容器,兼容空格问题
        containers = response.css('div.col-md-4.column')
        for container in containers:
            # 提取文本并去除空白,空值时返回默认字符串
            company_name = container.css('a::text').get(default='').strip()
            if company_name:
                yield {'name': company_name}

4. 添加User-Agent伪装(应对反爬)

如果运行后仍无法访问目标页面,可能是被网站反爬拦截,在settings.py中添加浏览器UA:

USER_AGENT = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/110.0.0.0 Safari/537.36'

验证方法

运行蜘蛛后,查看日志是否出现以下记录:

Crawled (200) <GET https://www.quicktransportsolutions.com/carrier/usa-trucking-companies.php>

如果出现这条日志,说明目标页面已成功访问,此时控制台会输出提取到的公司名称。

内容的提问来源于stack exchange,提问作者Miracle

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 09:05:14