You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Scrapy返回的数据赋值给变量?

如何将Scrapy爬取的数据直接存入变量而不是JSON文件

为什么你拿到的是<Deferred>对象

Scrapy基于Twisted异步框架,runner2.crawl()是异步执行的,它返回的Deferred对象代表一个尚未完成的异步任务,不是直接的爬取结果。你需要通过回调函数来获取任务完成后的实际数据。

解决方案1:在爬虫实例中存储数据+异步回调

直接在爬虫类里维护一个列表存储爬取到的item,然后利用Deferred的回调在爬虫完成后取出数据:

import scrapy
from scrapy.crawler import CrawlerRunner
from twisted.internet import reactor

class Livescores2(scrapy.Spider):
    name = 'Home'
    
    def __init__(self, *args, **kwargs):
        super().__init__(*args, **kwargs)
        self.scraped_data = []  # 初始化存储数据的列表

    def start_requests(self):
        yield scrapy.Request('https://www.livescores.com/football/turkey/super-lig/?tz=3&table=league-home')

    def parse(self, response):
        for total in response.css('td'):
            item = {'total': total.css('::text').get()}
            self.scraped_data.append(item)
            yield item  # 保留yield不影响其他逻辑,也可根据需要移除

# 爬虫完成后获取数据的回调函数
def fetch_scraped_data(crawler_result):
    return crawler_result.spider.scraped_data

# 处理最终数据的函数
def process_data(data):
    # 这里的data就是你要的实际爬取结果,可以赋值给变量或做后续处理
    print("实际爬取数据:")
    print(data)
    reactor.stop()  # 停止Twisted异步反应器

runner2 = CrawlerRunner()
deferred = runner2.crawl(Livescores2)
# 绑定回调,爬虫完成后先获取数据,再处理数据
deferred.addCallback(fetch_scraped_data)
deferred.addCallback(process_data)

# 启动异步任务
reactor.run()

解决方案2:用自定义Item Pipeline收集数据

通过Scrapy的Pipeline机制来统一收集所有item,适合更复杂的场景:

import scrapy
from scrapy.crawler import CrawlerRunner
from twisted.internet import reactor

# 自定义Pipeline用于收集数据
class DataCollectPipeline:
    def __init__(self):
        self.scraped_data = []

    def process_item(self, item, spider):
        self.scraped_data.append(item)
        return item

# 从Pipeline中取出数据的回调函数
def get_data_from_pipeline(crawler_result):
    # 获取注册的Pipeline实例
    pipeline = crawler_result.crawler.engine.scraper.item_pipelines[DataCollectPipeline]
    return pipeline.scraped_data

# 处理数据的函数
def handle_data(data):
    print("爬取到的数据:")
    print(data)
    reactor.stop()

# 配置Runner,注册自定义Pipeline
runner2 = CrawlerRunner(settings={
    'ITEM_PIPELINES': {
        '__main__.DataCollectPipeline': 300,  # 数字越小优先级越高
    }
})

deferred = runner2.crawl(Livescores2)
deferred.addCallback(get_data_from_pipeline)
deferred.addCallback(handle_data)

reactor.run()

注意事项

  • 必须启动Twisted的reactor.run()才能让异步任务执行完成,任务结束后要调用reactor.stop()终止反应器。
  • 如果你要在安卓APP中使用,需要考虑Python代码与安卓的集成方式(比如将Python服务作为后端,安卓通过HTTP请求获取数据,或者使用跨平台框架),但核心的数据收集逻辑可以参考上面的代码。

内容的提问来源于stack exchange,提问作者user19925544

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 18:40:30