使用BeautifulSoup爬取Hellopeter评论:无法获取核心数据求助
Hey there! I’ve run into this exact issue with SPA (Single Page Application) sites like Hellopeter before—let me break down what’s happening and how to fix it.
Why Your Current Code Isn’t Working
Hellopeter uses client-side JavaScript (likely a framework like Vue or React) to render content. When you use requests.get(), you’re only fetching the initial static HTML of the page, which is just the empty <div id="app"></div> shell. All those elements with data-v-* attributes (the rating, review body, and title) are generated dynamically by JavaScript after the page loads—so BeautifulSoup can’t see them in the static response.
Solution 1: Use a Headless Browser to Render JavaScript
The most reliable way to get the fully rendered DOM is to use a tool like Selenium or Playwright that simulates a real browser. These tools load the page, execute all JavaScript, and let you access the complete rendered content.
Example with Selenium:
First, install Selenium and download the appropriate browser driver (e.g., ChromeDriver for Chrome):
pip install selenium
Then use this code to scrape the elements you need:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup # Initialize Chrome in headless mode (no visible window) options = webdriver.ChromeOptions() options.add_argument("--headless=new") driver = webdriver.Chrome(options=options) try: # Load the target page url = 'https://www.hellopeter.com/telkom/reviews/appalling-service-cdc4559e8084e0db01dd6d7e807875460607ac77-2851593' driver.get(url) # Wait for the review content to render (adjust timeout if needed) wait = WebDriverWait(driver, 10) wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, 'p[itemprop="reviewBody"]'))) # Get the fully rendered page source page_source = driver.page_source soup = BeautifulSoup(page_source, 'html.parser') # Extract your target elements rating = soup.find('p', {'itemprop': 'reviewRating'}).get_text(strip=True) review_body = soup.find('p', {'itemprop': 'reviewBody'}).get_text(strip=True) review_title = soup.find('p', {'itemprop': 'name', 'class': 'is-detail-card'}).get_text(strip=True) print(f"Rating: {rating}") print(f"Title: {review_title}") print(f"Review: {review_body}") finally: # Make sure to close the browser driver.quit()
Solution 2: Directly Call the Site’s API (Faster Alternative)
Instead of scraping the DOM, you can often find the API endpoint that the site uses to fetch review data. Here’s how:
- Open your browser’s Developer Tools (F12) and go to the Network tab.
- Reload the Hellopeter review page.
- Look for XHR/fetch requests (filter by "XHR" or "Fetch")—you’ll likely see a request that returns JSON data containing the review details (rating, body, title, etc.).
- Copy that API URL and use
requeststo fetch the JSON directly. This is faster and more reliable than parsing HTML.
For example, if the API endpoint looks like https://www.hellopeter.com/api/reviews/2851593, you could do:
import requests headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } api_url = 'INSERT_THE_API_URL_HERE' response = requests.get(api_url, headers=headers) review_data = response.json() # Extract data directly from the JSON (adjust keys based on the actual response) rating = review_data['rating'] review_title = review_data['title'] review_body = review_data['body']
Important Notes
- Respect the Site’s Terms: Make sure Hellopeter allows scraping by checking their
robots.txtand Terms of Service. Avoid making too many rapid requests (add delays withtime.sleep()) to avoid getting blocked. - User-Agent: Always set a proper
User-Agentheader to mimic a real browser—this helps avoid being flagged as a bot.
内容的提问来源于stack exchange,提问作者ashrek

