You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy配置MongoDB Pipeline后无法存储爬取数据问题求助

问题排查及修复方案

1. 核心错误:ITEM_PIPELINES配置格式错误

你的settings.py中ITEM_PIPELINES使用了列表格式,Scrapy要求该配置必须为字典格式,键是pipeline的类路径,值是优先级(数字越小优先级越高),这是你的pipeline全程不生效、数据没法存入MongoDB的最主要原因。
修复后的settings.py对应部分:

ITEM_PIPELINES = {
    'realstatedata.pipelines.MongoPipeline': 300,
}

其余MONGO配置保持不变即可。

2. Pipeline代码兼容问题

你使用的insert()方法在PyMongo 3.0+版本已经被弃用,替换为insert_one()方法,否则会触发方法不存在的报错,数据无法写入。
修复后的pipeline.py的process_item方法:

def process_item(self, item, spider):
    self.db[self.collection_name].insert_one(dict(item))
    logging.debug("Properties added to MongoDB")
    return item

另外建议你先在本地测试MongoDB连接是否正常,27017端口是否开放、有无权限验证,如果有账号密码要在MONGO_URI里补上,格式为mongodb://用户名:密码@localhost:27017。

3. 数据提取逻辑错误

你当前的xpath提取规则存在重复取值、定位脆弱的问题:

  • prop_rooms和prop_bath两个字段使用了完全相同的xpath,最终两个字段都会取到第一个匹配的房间数值,卫浴数据根本拿不到
  • price_rent使用style属性定位,只要网站调整样式就会提取失效
  • 提取的文本没有做空白字符清理,存入数据库会有大量多余的换行、空格
  • 下一页判断逻辑错误,只要xpath能匹配到标签(不管有没有href值)都会触发翻页请求,容易出现无效请求
    修复后的spider代码:
import scrapy
from realstatedata.items import RealstatedataItem

class RsdataSpider(scrapy.Spider):
    name = 'realstatedata'
    allowed_domains = ['vivareal.com.br']
    start_urls = ['https://www.vivareal.com.br/aluguel/sp/sao-jose-dos-campos/apartamento_residencial/#preco-ate=2000']

    def parse(self, response):
        yield from self.scrape(response)
        # 修复下一页判断逻辑
        nextpage_path = response.xpath('//a[@title="Próxima página"]/@href').extract_first()
        if nextpage_path:
            path = '?' + nextpage_path[1:]
            nextpage = response.urljoin(path)
            print("Found url: {}".format(nextpage))
            yield scrapy.Request(nextpage)

    def scrape(self, response):
        for resource in response.xpath('//article[@class="property-card__container js-property-card"]/..'):
            item = RealstatedataItem()
            # 所有提取字段加strip处理,空值返回空字符串
            item['description'] = resource.xpath('.//h2/span[@class="property-card__title js-cardLink js-card-title"]/text()').extract_first(default='').strip()
            item['address'] = resource.xpath('.//span[@class="property-card__address"]/text()').extract_first(default='').strip()
            item['prop_area'] = resource.xpath('.//span[@class="property-card__detail-value js-property-card-value property-card__detail-area js-property-card-detail-area"]/text()').extract_first(default='').strip()
            # 修复房间、卫浴的xpath,按列表位序取值
            item['prop_rooms'] = resource.xpath('.//ul/li[2]/span[@class="property-card__detail-value js-property-card-value"]/text()').extract_first(default='').strip()
            item['prop_bath'] = resource.xpath('.//ul/li[3]/span[@class="property-card__detail-value js-property-card-value"]/text()').extract_first(default='').strip()
            item['prop_parking'] = resource.xpath('.//ul/li[4]/span[@class="property-card__detail-value js-property-card-value"]/text()').extract_first(default='').strip()
            # 替换租金的xpath定位规则,避免依赖style属性
            item['price_rent'] = resource.xpath('.//div[@class="property-card__price"]/p/text()').extract_first(default='').strip()
            item['price_cond'] = resource.xpath('.//strong[@class="js-condo-price"]/text()').extract_first(default='').strip()
            item['realstate_name'] = resource.xpath('.//picture/img/@alt').extract_first(default='').strip()
            yield item

验证步骤

修复完成后按以下顺序测试:

  1. 本地启动MongoDB服务,确认27017端口可正常连接
  2. 先运行scrapy crawl realstatedata -o test.csv验证提取逻辑是否正常,检查csv文件里有没有完整的字段数据
  3. 去掉输出参数直接运行爬虫,再用MongoDB客户端查看vivareal_db库的rent_properties集合里是否有数据写入

内容的提问来源于stack exchange,提问作者conradobio

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.07 08:18:02