如何用Python的BeautifulSoup与Requests爬取指定网站8页评论并合并至DataFrame
Solution to Scrape Multiple Pages of Reviews
Great question! To scrape up to 8 pages of reviews and combine them into a single DataFrame, we just need to add a loop that iterates through each page, constructs the correct URL for each page, and appends the reviews from each page to our list. Here's how to do it:
Step-by-Step Explanation
- Pagination URL Structure: The target site uses a
pagequery parameter to navigate between pages. For example, page 2's URL ishttps://www.mouthshut.com/product-reviews/Kotak-811-Mobile-Banking-reviews-925917218?page=2. - Loop Through Pages: We'll loop from page 1 to 8, building the URL for each page dynamically.
- Polite Scraping: Add a small delay between requests to avoid overwhelming the server (this helps prevent getting blocked).
- Collect Reviews: For each page, extract reviews the same way you did for the first page, and append them to the shared
reviewslist. - Create DataFrame: Once all pages are processed, convert the collected reviews into a DataFrame.
Full Modified Code
import requests from bs4 import BeautifulSoup import pandas as pd import time # For adding delays between requests reviews = [] # List to store all reviews across pages # Loop through pages 1 to 8 inclusive for page_num in range(1, 9): # Construct the URL for the current page url = f"https://www.mouthshut.com/product-reviews/Kotak-811-Mobile-Banking-reviews-925917218?page={page_num}" # Send GET request to the page response = requests.get(url) # Add a 2-second delay to be polite to the server time.sleep(2) # Parse the HTML content soup = BeautifulSoup(response.content, 'html5lib') # Extract all review containers on the current page review_divs = soup.select('div.col-10.review') # Iterate through each review container and extract data for element in review_divs: review_data = { 'Review_Title': element.a.text.strip(), 'URL': element.a['href'], 'Review': element.find('div', {'class': ['more', 'reviewdata']}).text.strip() } reviews.append(review_data) # Convert the collected reviews into a DataFrame df = pd.DataFrame(reviews) # Print the DataFrame (you can also save it to CSV with df.to_csv('reviews.csv', index=False)) print(df)
Important Notes
- Delay Adjustment: The
time.sleep(2)adds a 2-second pause between requests. You can adjust this (e.g., 1-3 seconds) depending on how strict the site's anti-scraping measures are. - Error Handling: For robustness, you might want to add try-except blocks to handle cases where a request fails (e.g., network issues, page not found). For example:
try: response = requests.get(url) response.raise_for_status() # Raise an error for HTTP status codes like 404 or 500 except requests.exceptions.RequestException as e: print(f"Error fetching page {page_num}: {e}") continue - Robust Selectors: Ensure that the CSS selectors (
div.col-10.reviewand the review text class) remain valid. Sites sometimes update their HTML structure, so you may need to adjust these if the scraping stops working.
内容的提问来源于stack exchange,提问作者Aayush Kaushal
相关产品推荐
相关产品推荐

