You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬虫爬取主页面及子页面时仅获取一条数据的问题求助

Hey there! Let's figure out why your Scrapy spider is only pulling one item instead of the 25 you're expecting. Here are the most common issues and fixes to check:

Common Causes & Solutions

1. Selector Mistakes (The #1 Culprit)

Chances are your selector is only targeting the first element instead of all matching ones. For example:

  • If you used extract_first() or get() directly on the parent selector (instead of looping through all matches), you'll only get the first item.
  • Or maybe your selector includes a hardcoded index like :nth-child(1) that restricts it to the first element.

Fix example:

# ❌ Wrong: Only extracts the first item
def parse(self, response):
    item = MyItem()
    item['title'] = response.css('h2.product-title::text').get()
    yield item

# ✅ Correct: Loops through all matching elements
def parse(self, response):
    # Grab all 25 item containers first
    product_containers = response.css('div.product-card')
    print(f"Found {len(product_containers)} items on page!") # Add this to debug
    
    for container in product_containers:
        item = MyItem()
        item['title'] = container.css('h2.product-title::text').get()
        item['price'] = container.css('span.price::text').get()
        yield item

2. Accidental Early Termination

Double-check your parse method for any break or return statements that stop the loop after the first item. It's easy to slip one in during testing!

Bad example to avoid:

def parse(self, response):
    for container in response.css('div.product-card'):
        item = MyItem()
        item['title'] = container.css('h2.product-title::text').get()
        yield item
        break # This kills the loop after the first item!

3. Duplicate Item Filtering

Scrapy automatically filters duplicate items by default. If your 25 items are being flagged as duplicates (maybe your Item class doesn't have unique identifiers), only the first one will be kept.

Test this:
Temporarily disable duplicate filtering in settings.py to see if all items show up:

DUPEFILTER_CLASS = 'scrapy.dupefilters.BaseDupeFilter'

Note: Re-enable this after testing—you don't want duplicate data in your final output.

4. Dynamic Content Loading

If the main page loads only 1 item initially and the rest load when scrolling, Scrapy's default HTML crawl won't see the other 24 (since they're loaded via JavaScript).

Check this:
Right-click the page → "View Page Source" and search for the 2nd item's title. If it's not there, you need to use a tool like scrapy-playwright or scrapy-splash to render the dynamic content.

5. Debug with Logs

Turn on debug logging to see what Scrapy is doing:

scrapy crawl your_spider_name -L DEBUG

Look for lines like Scraped from <200 ...> to count how many items are actually being scraped.

If none of these fix it, sharing a snippet of your parse method and the main page's HTML structure (for the item containers) would help narrow it down further!

内容的提问来源于stack exchange,提问作者user1796346

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 08:58:36