You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬虫使用runspider无法正常导出JSON文件问题咨询

Scrapy爬虫JSON导出异常问题解决

问题原因

你代码里多个字段直接返回了Scrapy的SelectorList对象(比如response.css('.product_main h1::text')),这类对象没法被JSON序列化,导致导出JSON时生成不了有效内容,最终出现只有逗号的情况。CSV能正常导出是因为Scrapy处理CSV时会自动提取SelectorList里的文本,但JSON序列化没这个自动处理逻辑。

解决方案

把所有返回SelectorList的字段,用.get()方法提取单个文本内容(如果要保留多个结果就用.getall()),确保每个字段都是JSON能识别的字符串类型。

修改后的代码

import scrapy
# from bookscraper.items import bootitem 
class BookspiderSpider(scrapy.Spider):
    name = "bookspider"
    allowed_domains = ["books.toscrape.com"]
    start_urls = ["https://books.toscrape.com/"]

    def parse(self, response):
        
        books = response.css('article.product_pod')

        for book in books:

            relative_url = book.css('h3 a').attrib['href']

            if 'catalogue/' in relative_url:
                book_url = 'https://books.toscrape.com/' + relative_url
            else:
                book_url = 'https://books.toscrape.com/catalogue/' + relative_url
            
            yield scrapy.Request(book_url, callback=self.parse_book_page)

        next_page = response.css('li.next a::attr(href)').get()
        
        if next_page is not None:
            if 'catalogue/' in next_page:
                next_page_url = 'https://books.toscrape.com/' + next_page
            else:
                next_page_url = 'https://books.toscrape.com/catalogue/' + next_page
            
            yield response.follow(next_page_url, callback=self.parse)

    def parse_book_page(self, response):
     
        table_rows = response.css('table tr')
       #scra book_item = bookitem()
        yield{
                'url' : response.url,
                'title': response.css('.product_main h1::text').get(),
                'product_type': table_rows[1].css('td::text').get(),
                'price_excl_tax': table_rows[2].css('td::text').get(),
                'price_incl_tax': table_rows[3].css('td::text').get(),
                'tax' : table_rows[4].css('td::text').get(),
                'availability' : table_rows[5].css('td::text').get(),
                'num_reviews': table_rows[6].css('td::text').get(),
                'stars': response.css('p.star-rating').attrib['class'],
                'category': response.xpath("//ul[@class='breadcrumb']/li[@class='active']/preceding-sibling::li[1]/a/text()").get(),
                'description': response.xpath("//div[@id='product_description']/following-sibling::p/text()").get(),
                'price': response.css('p.price-color::text').get(),
        }

补充说明

  • 要是某个字段可能有多个结果需要全部保留,可以用.getall()替代.get(),会返回字符串列表,JSON也能正常序列化。
  • 这个问题和runspider语法无关,不管用crawl还是runspider,只要字段是可序列化类型,都能正常导出JSON。

内容的提问来源于stack exchange,提问作者umer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 17:55:11