You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy仅返回首条数据求助:如何获取所有条目信息

Fixing Your Scrapy Spider: Returning All Items Instead of Just the First

Hey there! I see exactly what's tripping up your spider—let's break it down and fix it step by step.

The Core Issue

Your current loop is iterating over the entire <ul class="card-collection"> container instead of each individual list item (<li class="card-item">) inside it. That's why you only get the first result: you're processing one container, not each listing within it.

On top of that, you're using response.xpath() inside the loop instead of row.xpath(), which means you're always querying the full page instead of the specific item you're looping over.

Corrected Code

Here's the fixed version of your spider, with comments explaining the key changes:

import scrapy
class FarmtoolsSpider(scrapy.Spider):
    name = 'farmtools'
    allowed_domains = ['www.donedeal.ie']
    start_urls = ['https://www.donedeal.ie/farmtools/']
    
    def parse(self, response):
        # Loop over each individual listing (li element) inside the collection
        for row in response.xpath('//ul[@class="card-collection"]/li[@class="card-item"]'):
            yield {
                # Use row.xpath() to target only the current item's elements
                'item_title': row.xpath('.//p[@class="card__body-title"]/text()').get().strip(),
                'item_county': row.xpath('.//ul[@class="card__body-keyinfo"]/li[2]/text()').get().strip(),
                'item_price': row.xpath('.//p[@class="card__price"]/span[1]/text()').get().strip(),
                'item_id': row.xpath('./a/@href').get().split('/')[-1]  # Extract clean ID from the URL
            }

What Changed (And Why)

  • Target individual listings: The loop now focuses on each <li class="card-item"> inside the collection—each row represents one farm tool entry.
  • Use row.xpath(): This ensures you only pull data from the current item in the loop, not the entire page.
  • .strip(): Added to clean up messy whitespace in the extracted text.
  • Cleaner item ID: Split the href URL to grab just the unique ID at the end (optional, but makes the ID more usable).

Why getall() Gave You "Blocked" Results

When you used getall(), it pulled every matching element from the entire page at once, which is why you got a jumbled chunk of data instead of structured individual entries. By looping over each item first, you can extract neat, separate data for every listing.

Give this a run, and you should get every farm tool entry with its title, county, price, and ID!

内容的提问来源于stack exchange,提问作者billiam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 12:17:43