Scrapy仅返回首条数据求助:如何获取所有条目信息
Hey there! I see exactly what's tripping up your spider—let's break it down and fix it step by step.
The Core Issue
Your current loop is iterating over the entire <ul class="card-collection"> container instead of each individual list item (<li class="card-item">) inside it. That's why you only get the first result: you're processing one container, not each listing within it.
On top of that, you're using response.xpath() inside the loop instead of row.xpath(), which means you're always querying the full page instead of the specific item you're looping over.
Corrected Code
Here's the fixed version of your spider, with comments explaining the key changes:
import scrapy class FarmtoolsSpider(scrapy.Spider): name = 'farmtools' allowed_domains = ['www.donedeal.ie'] start_urls = ['https://www.donedeal.ie/farmtools/'] def parse(self, response): # Loop over each individual listing (li element) inside the collection for row in response.xpath('//ul[@class="card-collection"]/li[@class="card-item"]'): yield { # Use row.xpath() to target only the current item's elements 'item_title': row.xpath('.//p[@class="card__body-title"]/text()').get().strip(), 'item_county': row.xpath('.//ul[@class="card__body-keyinfo"]/li[2]/text()').get().strip(), 'item_price': row.xpath('.//p[@class="card__price"]/span[1]/text()').get().strip(), 'item_id': row.xpath('./a/@href').get().split('/')[-1] # Extract clean ID from the URL }
What Changed (And Why)
- Target individual listings: The loop now focuses on each
<li class="card-item">inside the collection—eachrowrepresents one farm tool entry. - Use
row.xpath(): This ensures you only pull data from the current item in the loop, not the entire page. .strip(): Added to clean up messy whitespace in the extracted text.- Cleaner item ID: Split the href URL to grab just the unique ID at the end (optional, but makes the ID more usable).
Why getall() Gave You "Blocked" Results
When you used getall(), it pulled every matching element from the entire page at once, which is why you got a jumbled chunk of data instead of structured individual entries. By looping over each item first, you can extract neat, separate data for every listing.
Give this a run, and you should get every farm tool entry with its title, county, price, and ID!
内容的提问来源于stack exchange,提问作者billiam

