Python爬取Google Play评论失败求助:返回空列表问题
Hey there! I totally get the frustration when your scraper spits out an empty list—been there, done that. Let’s break down why this is happening and fix it step by step.
Why Your Current Code Isn’t Working
Google Play loads reviews dynamically using JavaScript. Tools like urllib only fetch the initial static HTML sent by the server, which doesn’t include the reviews yet—those get loaded later once the browser runs the page’s JS. That’s why BeautifulSoup can’t find the UD7Dzf or other review-related nodes you’re targeting: they aren’t in the initial HTML at all!
The Solution: Use a Headless Browser to Render Dynamic Content
To grab the fully loaded page (with reviews included), we’ll use selenium—a tool that simulates a real browser, runs the JavaScript, and lets us scrape the fully rendered page. Here’s how to make it work:
Step 1: Install Required Tools
First, install selenium via pip:
pip install selenium
You’ll also need to download a browser driver (e.g., ChromeDriver for Chrome) that matches your browser version. Place the driver executable in a location your Python script can access, or add it to your system PATH.
Step 2: Working Scraper Code
This script will simulate scrolling to load more reviews, then extract the text:
from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.chrome.options import Options import time # Configure Chrome to run in headless mode (no visible browser window) chrome_options = Options() chrome_options.add_argument("--headless=new") chrome_options.add_argument("--disable-gpu") # Initialize the browser driver = webdriver.Chrome(options=chrome_options) # Get the Google Play reviews page (make sure it's the reviews tab URL!) url = input("Enter the Google Play app reviews URL: ") driver.get(url) # Scroll to load all reviews (Google Play loads more as you scroll) scroll_pause = 2 last_scroll_height = driver.execute_script("return document.body.scrollHeight") while True: # Scroll to the bottom of the page driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") # Wait for new reviews to load time.sleep(scroll_pause) # Check if we've reached the end of the page new_scroll_height = driver.execute_script("return document.body.scrollHeight") if new_scroll_height == last_scroll_height: break last_scroll_height = new_scroll_height # Grab the fully rendered page source page_html = driver.page_source driver.quit() # Close the browser # Parse with BeautifulSoup soup = BeautifulSoup(page_html, "html.parser") reviews_list = [] # Extract review text (target the correct span class) for review_element in soup.find_all("span", class_="UD7Dzf"): review_text = review_element.get_text(strip=True) if review_text: # Skip empty entries reviews_list.append(review_text) # Print the results print(f"Found {len(reviews_list)} reviews:") for i, review in enumerate(reviews_list, 1): print(f"{i}. {review}")
Key Notes:
- Page Structure Changes: Google Play might update its HTML classes/structure over time. If this stops working, right-click a review in your browser, select "Inspect", and find the new CSS selector for the review text.
- Rate Limiting: Don’t scrape too aggressively—Google might block your IP if you make too many requests too quickly. Add longer pauses if needed.
- Alternative: Official API: For a more reliable solution, consider using the Google Play Developer API (you’ll need a Google Developer account and API credentials). This is the best way to get reviews without worrying about web scraping blocks or structure changes.
内容的提问来源于stack exchange,提问作者user13215247

