You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取Hellopeter评论:无法获取核心数据求助

How to Scrape Dynamic Hellopeter Reviews for Sentiment Analysis

Hey there! I’ve run into this exact issue with SPA (Single Page Application) sites like Hellopeter before—let me break down what’s happening and how to fix it.

Why Your Current Code Isn’t Working

Hellopeter uses client-side JavaScript (likely a framework like Vue or React) to render content. When you use requests.get(), you’re only fetching the initial static HTML of the page, which is just the empty <div id="app"></div> shell. All those elements with data-v-* attributes (the rating, review body, and title) are generated dynamically by JavaScript after the page loads—so BeautifulSoup can’t see them in the static response.

Solution 1: Use a Headless Browser to Render JavaScript

The most reliable way to get the fully rendered DOM is to use a tool like Selenium or Playwright that simulates a real browser. These tools load the page, execute all JavaScript, and let you access the complete rendered content.

Example with Selenium:

First, install Selenium and download the appropriate browser driver (e.g., ChromeDriver for Chrome):

pip install selenium

Then use this code to scrape the elements you need:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup

# Initialize Chrome in headless mode (no visible window)
options = webdriver.ChromeOptions()
options.add_argument("--headless=new")
driver = webdriver.Chrome(options=options)

try:
    # Load the target page
    url = 'https://www.hellopeter.com/telkom/reviews/appalling-service-cdc4559e8084e0db01dd6d7e807875460607ac77-2851593'
    driver.get(url)

    # Wait for the review content to render (adjust timeout if needed)
    wait = WebDriverWait(driver, 10)
    wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, 'p[itemprop="reviewBody"]')))

    # Get the fully rendered page source
    page_source = driver.page_source
    soup = BeautifulSoup(page_source, 'html.parser')

    # Extract your target elements
    rating = soup.find('p', {'itemprop': 'reviewRating'}).get_text(strip=True)
    review_body = soup.find('p', {'itemprop': 'reviewBody'}).get_text(strip=True)
    review_title = soup.find('p', {'itemprop': 'name', 'class': 'is-detail-card'}).get_text(strip=True)

    print(f"Rating: {rating}")
    print(f"Title: {review_title}")
    print(f"Review: {review_body}")

finally:
    # Make sure to close the browser
    driver.quit()

Solution 2: Directly Call the Site’s API (Faster Alternative)

Instead of scraping the DOM, you can often find the API endpoint that the site uses to fetch review data. Here’s how:

  1. Open your browser’s Developer Tools (F12) and go to the Network tab.
  2. Reload the Hellopeter review page.
  3. Look for XHR/fetch requests (filter by "XHR" or "Fetch")—you’ll likely see a request that returns JSON data containing the review details (rating, body, title, etc.).
  4. Copy that API URL and use requests to fetch the JSON directly. This is faster and more reliable than parsing HTML.

For example, if the API endpoint looks like https://www.hellopeter.com/api/reviews/2851593, you could do:

import requests

headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
}
api_url = 'INSERT_THE_API_URL_HERE'
response = requests.get(api_url, headers=headers)
review_data = response.json()

# Extract data directly from the JSON (adjust keys based on the actual response)
rating = review_data['rating']
review_title = review_data['title']
review_body = review_data['body']

Important Notes

  • Respect the Site’s Terms: Make sure Hellopeter allows scraping by checking their robots.txt and Terms of Service. Avoid making too many rapid requests (add delays with time.sleep()) to avoid getting blocked.
  • User-Agent: Always set a proper User-Agent header to mimic a real browser—this helps avoid being flagged as a bot.

内容的提问来源于stack exchange,提问作者ashrek

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 07:40:31