You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页数据爬取异常:仅获取10页数据,目标页面共26页

Fixing Flipkart Scraping: Only Getting 10 Pages Instead of 26

Hey there! Let's break down why you're only pulling 10 pages of laptop listings instead of the full 26, and how to fix it.

Common Issues & Solutions

1. Anti-Scraping Blocks (Most Likely Culprit)

Flipkart actively blocks non-browser requests to prevent scraping. Your current code sends requests without any browser-like headers, so the site is limiting you to 10 pages as a protective measure.

Fix: Add a User-Agent header to mimic a real browser, and add small delays between requests to avoid triggering rate limits. You can also include other headers like Accept-Language for more authenticity.

2. Incorrect Page Number Extraction

Your code grabs the last page number from pagination links with:

page_nr=soup.find_all("a",{"class":"_33m_Yg"})[-1].text

But Flipkart truncates pagination links for large page counts (e.g., showing 1...10 11...26 instead of all 26 numbers). So you're only getting the last visible page number (10) instead of the actual total (26).

Fix: Either:

  • Loop directly from 1 to 26 if you know the total pages upfront
  • Check the page source for hidden elements/script tags that contain the total page count (look for keywords like totalPages)
  • Keep looping until the returned page has no product listings (indicating you've passed the last valid page)

Flipkart sometimes requires maintaining a session to access beyond a certain number of pages. Using a requests.Session() instead of individual get calls preserves cookies across requests, which can help bypass session-based restrictions.

Corrected Code Example

import requests
from bs4 import BeautifulSoup
import time

# Initialize a session to persist cookies
session = requests.Session()

# Add browser-like headers to avoid being flagged
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36",
    "Accept-Language": "en-US,en;q=0.9"
}

base_url = "https://www.flipkart.com/search?as=on&as-pos=1_1_ic_lapto&as-show=on&otracker=start&page={}&q=laptop&sid=6bo%2Fb5g&viewType=list"
product_list = []

# Loop through all 26 pages (adjust to dynamic detection if needed)
for page in range(1, 27):
    print(f"Scraping page {page}...")
    url = base_url.format(page)
    
    try:
        # Send request with headers and session
        response = session.get(url, headers=headers)
        response.raise_for_status()  # Catch HTTP errors like 403/404
        soup = BeautifulSoup(response.content, "html.parser")
        
        # Fetch product listings
        products = soup.find_all("div", {"class": "col _2-gKeQ"})
        
        # Stop loop if no products are found (end of listings)
        if not products:
            print(f"No products found on page {page}. Stopping early.")
            break
        
        # Extract data from each product
        for product in products:
            data = {}
            # Extract price (example field)
            price = product.find("div", {"class": "_1vC4OE _2rQ-NK"})
            data["price"] = price.text if price else "N/A"
            # Add other fields like product name, rating here
            product_list.append(data)
        
        # Add a 2-second delay to avoid rate limiting
        time.sleep(2)
    
    except Exception as e:
        print(f"Error scraping page {page}: {str(e)}")
        continue

print(f"Scraped total {len(product_list)} products from {page} pages.")

Key Improvements in This Code

  • Uses requests.Session() to maintain session state and cookies
  • Includes realistic headers to bypass basic anti-scraping checks
  • Adds delays between requests to avoid being blocked
  • Handles errors gracefully and stops early if no products are found
  • Explicitly loops through the known 26 pages (easily adjustable for dynamic page count detection)

内容的提问来源于stack exchange,提问作者Shaelander Chauhan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:21:28