Scrapy爬虫爬取主页面及子页面时仅获取一条数据的问题求助
Hey there! Let's figure out why your Scrapy spider is only pulling one item instead of the 25 you're expecting. Here are the most common issues and fixes to check:
1. Selector Mistakes (The #1 Culprit)
Chances are your selector is only targeting the first element instead of all matching ones. For example:
- If you used
extract_first()orget()directly on the parent selector (instead of looping through all matches), you'll only get the first item. - Or maybe your selector includes a hardcoded index like
:nth-child(1)that restricts it to the first element.
Fix example:
# ❌ Wrong: Only extracts the first item def parse(self, response): item = MyItem() item['title'] = response.css('h2.product-title::text').get() yield item # ✅ Correct: Loops through all matching elements def parse(self, response): # Grab all 25 item containers first product_containers = response.css('div.product-card') print(f"Found {len(product_containers)} items on page!") # Add this to debug for container in product_containers: item = MyItem() item['title'] = container.css('h2.product-title::text').get() item['price'] = container.css('span.price::text').get() yield item
2. Accidental Early Termination
Double-check your parse method for any break or return statements that stop the loop after the first item. It's easy to slip one in during testing!
Bad example to avoid:
def parse(self, response): for container in response.css('div.product-card'): item = MyItem() item['title'] = container.css('h2.product-title::text').get() yield item break # This kills the loop after the first item!
3. Duplicate Item Filtering
Scrapy automatically filters duplicate items by default. If your 25 items are being flagged as duplicates (maybe your Item class doesn't have unique identifiers), only the first one will be kept.
Test this:
Temporarily disable duplicate filtering in settings.py to see if all items show up:
DUPEFILTER_CLASS = 'scrapy.dupefilters.BaseDupeFilter'
Note: Re-enable this after testing—you don't want duplicate data in your final output.
4. Dynamic Content Loading
If the main page loads only 1 item initially and the rest load when scrolling, Scrapy's default HTML crawl won't see the other 24 (since they're loaded via JavaScript).
Check this:
Right-click the page → "View Page Source" and search for the 2nd item's title. If it's not there, you need to use a tool like scrapy-playwright or scrapy-splash to render the dynamic content.
5. Debug with Logs
Turn on debug logging to see what Scrapy is doing:
scrapy crawl your_spider_name -L DEBUG
Look for lines like Scraped from <200 ...> to count how many items are actually being scraped.
If none of these fix it, sharing a snippet of your parse method and the main page's HTML structure (for the item containers) would help narrow it down further!
内容的提问来源于stack exchange,提问作者user1796346

