Scrapy爬虫在Python脚本中返回重复最后条目问题排查
Alright, let's dig into this issue—this is a super common pitfall when calling Scrapy from a Python script, and it's almost always related to how you're creating your Item objects in your spider.
Root Cause
The problem happens when you reuse the same Item (or dictionary) instance across all your scraped items instead of creating a new one for each entry.
Here's the breakdown:
- If you define your Item object once (e.g., at the top of your Spider class, or outside the
parsemethod), every time you update its fields andyieldit, you're actually yielding a reference to the same object. - By the time your script collects all results in a list, all those references point to the final state of that single Item object—hence why you only see the last scraped entry repeated.
The scrapy crawl command works correctly because Scrapy's command-line pipeline implicitly handles serialization of Items as they're yielded, effectively creating a copy of the Item's state at the time of yielding. When you run Scrapy via a script, you're directly collecting object references, so this implicit copy doesn't happen.
How to Verify This
Add a quick debug print in your spider's parse method to check the unique ID of the Item object each time you yield it:
def parse(self, response): # ... your scraping logic ... item = MyItem() # Or whatever your Item class is item["title"] = response.css("h1::text").get() # Print the object ID to confirm uniqueness print(f"Yielding Item with ID: {id(item)}") yield item
If every printed ID is identical, you're definitely reusing the same Item instance.
Fix: Create a New Item Instance for Each Entry
The fix is straightforward: move the Item instantiation inside your parse method (or inside the loop where you process multiple items per response).
Example of Wrong Code
class MySpider(scrapy.Spider): name = "my_spider" start_urls = ["https://example.com"] # ❌ Item defined once, reused for all entries item = MyItem() def parse(self, response): for product in response.css(".product"): self.item["name"] = product.css(".name::text").get() self.item["price"] = product.css(".price::text").get() yield self.item
Example of Fixed Code
class MySpider(scrapy.Spider): name = "my_spider" start_urls = ["https://example.com"] def parse(self, response): for product in response.css(".product"): # ✅ New Item instance created for each product item = MyItem() item["name"] = product.css(".name::text").get() item["price"] = product.css(".price::text").get() yield item
This ensures each yield sends a unique Item object with its own state, so when your script collects the results, you'll get all distinct entries.
What If You're Using Dictionaries Instead of Scrapy Items?
Same rule applies—don't reuse the same dictionary. Create a new dict() for each entry:
# Wrong product_dict = {} for product in response.css(".product"): product_dict["name"] = ... yield product_dict # Right for product in response.css(".product"): product_dict = {} product_dict["name"] = ... yield product_dict
Even if you're using CrawlerRunner, this fix will resolve the duplicate entry issue because the core problem lies in the spider's Item creation logic, not the controller script.
内容的提问来源于stack exchange,提问作者Chris1309

