You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python的BeautifulSoup与Requests爬取指定网站8页评论并合并至DataFrame

Solution to Scrape Multiple Pages of Reviews

Great question! To scrape up to 8 pages of reviews and combine them into a single DataFrame, we just need to add a loop that iterates through each page, constructs the correct URL for each page, and appends the reviews from each page to our list. Here's how to do it:

Step-by-Step Explanation

  1. Pagination URL Structure: The target site uses a page query parameter to navigate between pages. For example, page 2's URL is https://www.mouthshut.com/product-reviews/Kotak-811-Mobile-Banking-reviews-925917218?page=2.
  2. Loop Through Pages: We'll loop from page 1 to 8, building the URL for each page dynamically.
  3. Polite Scraping: Add a small delay between requests to avoid overwhelming the server (this helps prevent getting blocked).
  4. Collect Reviews: For each page, extract reviews the same way you did for the first page, and append them to the shared reviews list.
  5. Create DataFrame: Once all pages are processed, convert the collected reviews into a DataFrame.

Full Modified Code

import requests
from bs4 import BeautifulSoup
import pandas as pd
import time  # For adding delays between requests

reviews = []  # List to store all reviews across pages

# Loop through pages 1 to 8 inclusive
for page_num in range(1, 9):
    # Construct the URL for the current page
    url = f"https://www.mouthshut.com/product-reviews/Kotak-811-Mobile-Banking-reviews-925917218?page={page_num}"
    
    # Send GET request to the page
    response = requests.get(url)
    
    # Add a 2-second delay to be polite to the server
    time.sleep(2)
    
    # Parse the HTML content
    soup = BeautifulSoup(response.content, 'html5lib')
    
    # Extract all review containers on the current page
    review_divs = soup.select('div.col-10.review')
    
    # Iterate through each review container and extract data
    for element in review_divs:
        review_data = {
            'Review_Title': element.a.text.strip(),
            'URL': element.a['href'],
            'Review': element.find('div', {'class': ['more', 'reviewdata']}).text.strip()
        }
        reviews.append(review_data)

# Convert the collected reviews into a DataFrame
df = pd.DataFrame(reviews)

# Print the DataFrame (you can also save it to CSV with df.to_csv('reviews.csv', index=False))
print(df)

Important Notes

  • Delay Adjustment: The time.sleep(2) adds a 2-second pause between requests. You can adjust this (e.g., 1-3 seconds) depending on how strict the site's anti-scraping measures are.
  • Error Handling: For robustness, you might want to add try-except blocks to handle cases where a request fails (e.g., network issues, page not found). For example:
    try:
        response = requests.get(url)
        response.raise_for_status()  # Raise an error for HTTP status codes like 404 or 500
    except requests.exceptions.RequestException as e:
        print(f"Error fetching page {page_num}: {e}")
        continue
    
  • Robust Selectors: Ensure that the CSS selectors (div.col-10.review and the review text class) remain valid. Sites sometimes update their HTML structure, so you may need to adjust these if the scraping stops working.

内容的提问来源于stack exchange,提问作者Aayush Kaushal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 07:43:00