You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬虫在Python脚本中返回重复最后条目问题排查

Alright, let's dig into this issue—this is a super common pitfall when calling Scrapy from a Python script, and it's almost always related to how you're creating your Item objects in your spider.

Root Cause

The problem happens when you reuse the same Item (or dictionary) instance across all your scraped items instead of creating a new one for each entry.

Here's the breakdown:

  • If you define your Item object once (e.g., at the top of your Spider class, or outside the parse method), every time you update its fields and yield it, you're actually yielding a reference to the same object.
  • By the time your script collects all results in a list, all those references point to the final state of that single Item object—hence why you only see the last scraped entry repeated.

The scrapy crawl command works correctly because Scrapy's command-line pipeline implicitly handles serialization of Items as they're yielded, effectively creating a copy of the Item's state at the time of yielding. When you run Scrapy via a script, you're directly collecting object references, so this implicit copy doesn't happen.

How to Verify This

Add a quick debug print in your spider's parse method to check the unique ID of the Item object each time you yield it:

def parse(self, response):
    # ... your scraping logic ...
    item = MyItem()  # Or whatever your Item class is
    item["title"] = response.css("h1::text").get()
    # Print the object ID to confirm uniqueness
    print(f"Yielding Item with ID: {id(item)}")
    yield item

If every printed ID is identical, you're definitely reusing the same Item instance.

Fix: Create a New Item Instance for Each Entry

The fix is straightforward: move the Item instantiation inside your parse method (or inside the loop where you process multiple items per response).

Example of Wrong Code

class MySpider(scrapy.Spider):
    name = "my_spider"
    start_urls = ["https://example.com"]
    
    # ❌ Item defined once, reused for all entries
    item = MyItem()

    def parse(self, response):
        for product in response.css(".product"):
            self.item["name"] = product.css(".name::text").get()
            self.item["price"] = product.css(".price::text").get()
            yield self.item

Example of Fixed Code

class MySpider(scrapy.Spider):
    name = "my_spider"
    start_urls = ["https://example.com"]

    def parse(self, response):
        for product in response.css(".product"):
            # ✅ New Item instance created for each product
            item = MyItem()
            item["name"] = product.css(".name::text").get()
            item["price"] = product.css(".price::text").get()
            yield item

This ensures each yield sends a unique Item object with its own state, so when your script collects the results, you'll get all distinct entries.

What If You're Using Dictionaries Instead of Scrapy Items?

Same rule applies—don't reuse the same dictionary. Create a new dict() for each entry:

# Wrong
product_dict = {}
for product in response.css(".product"):
    product_dict["name"] = ...
    yield product_dict

# Right
for product in response.css(".product"):
    product_dict = {}
    product_dict["name"] = ...
    yield product_dict

Even if you're using CrawlerRunner, this fix will resolve the duplicate entry issue because the core problem lies in the spider's Item creation logic, not the controller script.

内容的提问来源于stack exchange,提问作者Chris1309

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 08:36:07