网页数据爬取异常:仅获取10页数据,目标页面共26页
Hey there! Let's break down why you're only pulling 10 pages of laptop listings instead of the full 26, and how to fix it.
Common Issues & Solutions
1. Anti-Scraping Blocks (Most Likely Culprit)
Flipkart actively blocks non-browser requests to prevent scraping. Your current code sends requests without any browser-like headers, so the site is limiting you to 10 pages as a protective measure.
Fix: Add a User-Agent header to mimic a real browser, and add small delays between requests to avoid triggering rate limits. You can also include other headers like Accept-Language for more authenticity.
2. Incorrect Page Number Extraction
Your code grabs the last page number from pagination links with:
page_nr=soup.find_all("a",{"class":"_33m_Yg"})[-1].text
But Flipkart truncates pagination links for large page counts (e.g., showing 1...10 11...26 instead of all 26 numbers). So you're only getting the last visible page number (10) instead of the actual total (26).
Fix: Either:
- Loop directly from 1 to 26 if you know the total pages upfront
- Check the page source for hidden elements/script tags that contain the total page count (look for keywords like
totalPages) - Keep looping until the returned page has no product listings (indicating you've passed the last valid page)
3. Session & Cookie Issues
Flipkart sometimes requires maintaining a session to access beyond a certain number of pages. Using a requests.Session() instead of individual get calls preserves cookies across requests, which can help bypass session-based restrictions.
Corrected Code Example
import requests from bs4 import BeautifulSoup import time # Initialize a session to persist cookies session = requests.Session() # Add browser-like headers to avoid being flagged headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36", "Accept-Language": "en-US,en;q=0.9" } base_url = "https://www.flipkart.com/search?as=on&as-pos=1_1_ic_lapto&as-show=on&otracker=start&page={}&q=laptop&sid=6bo%2Fb5g&viewType=list" product_list = [] # Loop through all 26 pages (adjust to dynamic detection if needed) for page in range(1, 27): print(f"Scraping page {page}...") url = base_url.format(page) try: # Send request with headers and session response = session.get(url, headers=headers) response.raise_for_status() # Catch HTTP errors like 403/404 soup = BeautifulSoup(response.content, "html.parser") # Fetch product listings products = soup.find_all("div", {"class": "col _2-gKeQ"}) # Stop loop if no products are found (end of listings) if not products: print(f"No products found on page {page}. Stopping early.") break # Extract data from each product for product in products: data = {} # Extract price (example field) price = product.find("div", {"class": "_1vC4OE _2rQ-NK"}) data["price"] = price.text if price else "N/A" # Add other fields like product name, rating here product_list.append(data) # Add a 2-second delay to avoid rate limiting time.sleep(2) except Exception as e: print(f"Error scraping page {page}: {str(e)}") continue print(f"Scraped total {len(product_list)} products from {page} pages.")
Key Improvements in This Code
- Uses
requests.Session()to maintain session state and cookies - Includes realistic headers to bypass basic anti-scraping checks
- Adds delays between requests to avoid being blocked
- Handles errors gracefully and stops early if no products are found
- Explicitly loops through the known 26 pages (easily adjustable for dynamic page count detection)
内容的提问来源于stack exchange,提问作者Shaelander Chauhan

