You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Scrapy爬取网页无预期输出且JSON文件为空问题求助

搞定Scrapy爬取无输出&JSON空文件的问题

Hey,我来帮你捋捋这两个头疼的问题——爬不到预期内容、JSON文件为空,大概率是代码里几个关键环节没做好,咱们一步步拆解:

1. 先解决「爬两个网页」的问题

你说要爬两个网页,但目前start_urls里只放了一个链接,分两种情况处理:

  • 如果两个是独立的列表页,直接把第二个URL加到start_urls里就行:
    start_urls = [
        "http://www.greatplantpicks.org/plantlists/by_plant_type/conifer",
        "http://www.greatplantpicks.org/你要爬的第二个页面链接"
    ]
    
  • 如果第二个页面是第一个页面里的详情页(比如点进植物卡片看详情),那得在parse里先抓详情页链接,再发起请求:
    def parse(self, response):
        # 先抓页面里的详情页链接(这里的选择器要根据实际页面结构修改)
        detail_links = response.css('a.plant-detail::attr(href)').getall()
        for link in detail_links:
            # 拼接完整URL,避免相对路径出错
            full_link = response.urljoin(link)
            # 把请求交给专门处理详情页的函数
            yield scrapy.Request(full_link, callback=self.parse_detail)
    
    # 新增处理详情页的函数
    def parse_detail(self, response):
        item = PlantsItem()
        # 这里填详情页的数据提取逻辑
        item['plant_name'] = response.css('h1::text').get().strip()
        # 其他字段...
        yield item
    

2. JSON空文件的核心原因:没yield Item!

这是新手常踩的坑——Scrapy只有收到你yield出来的PlantsItem对象,才会把数据写入输出文件。你现在的parse方法是空的,自然啥也存不进去。

给你补个完整的示例,针对你那个针叶树列表页提取数据:

from plants.items import PlantsItem
import scrapy

class PlantsSpider(scrapy.Spider):
    name = "plants"
    allowed_domains = ["greatplantpicks.org"]
    # 这里放两个要爬的URL
    start_urls = [
        "http://www.greatplantpicks.org/plantlists/by_plant_type/conifer",
        "http://www.greatplantpicks.org/plantlists/by_plant_type/xxxx"
    ]

    def parse(self, response):
        # 先找页面上的所有植物条目(选择器要根据实际页面结构调整!)
        plant_cards = response.css('div.plant-item')
        print(f"找到{len(plant_cards)}个植物条目")  # 先打印数量,确认能抓到内容

        for card in plant_cards:
            item = PlantsItem()
            # 提取常用名和学名,示例选择器,你要自己去浏览器里查元素改
            item['common_name'] = card.css('h3.common::text').get(default='').strip()
            item['scientific_name'] = card.css('p.sci-name::text').get(default='').strip()
            # 其他需要的字段都这么提取
            yield item  # 关键!必须yield,不然Scrapy收不到数据

        # 如果页面有分页,还可以抓下一页继续爬
        next_page = response.css('a.next::attr(href)').get()
        if next_page:
            yield response.follow(next_page, callback=self.parse)

3. 验证和调试技巧

  • 运行爬虫时,用这个命令导出JSON:
    scrapy crawl plants -o plants.json
    
  • 如果还是空的,先加print调试:比如打印当前响应的URL,打印抓到的条目数量,确认爬虫确实访问到了页面,且选择器能抓到内容。
  • 要是print没输出,大概率是CSS选择器写错了——打开浏览器开发者工具,对着要提取的元素右键「检查」,复制正确的选择器。

最后补个小细节

确保你的items.py里定义了对应的字段,不然提取的数据存不进Item:

import scrapy

class PlantsItem(scrapy.Item):
    common_name = scrapy.Field()
    scientific_name = scrapy.Field()
    # 其他需要的字段都在这定义

内容的提问来源于stack exchange,提问作者M.liza

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 10:32:39