You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Rvest爬取Steam评论的两类技术问题求助

Hey there, let's tackle your two scraping issues step by step—they're super common when dealing with dynamic sites like Steam, so you're in good company!

1. Separating Review Content from Post Date

Looking at your code, the review variable pulls the entire text from div.apphub_CardTextContent, which includes both the post date (the small top line) and the actual review body. Since you're already extracting the date separately, we can just strip that date string from the raw review text to get clean content.

Here's how to adjust your code with the stringr package to handle this:

library(rvest)
library(stringr)

link <- "https://steamcommunity.com/app/215080/reviews/"
page <- read_html(link) # Only read the page once for efficiency!

# Extract all fields
date <- page %>% 
  html_nodes("div.date_posted") %>% 
  html_text() %>% 
  str_trim() # Clean up extra whitespace

raw_review <- page %>% 
  html_nodes("div.apphub_CardTextContent") %>% 
  html_text()

# Remove the matching date from each review and trim leftover whitespace
review <- map2_chr(raw_review, date, ~str_remove_all(.x, .y) %>% str_trim())

opinion <- page %>% 
  html_nodes("div.title") %>% 
  html_text() %>% 
  str_trim()

hoursplayed <- page %>% 
  html_nodes("div.hours") %>% 
  html_text() %>% 
  str_trim()

helpful <- page %>% 
  html_nodes("div.found_helpful") %>% 
  html_text() %>% 
  str_trim()

# Build your data frame
tab <- data.frame(
  "Posted" = date,
  "Review" = review,
  "Opinion" = opinion,
  "Hours Played" = hoursplayed,
  "Number of helpful vote" = helpful
)

A quick note: I adjusted your code to read the page only once instead of multiple times—this saves bandwidth and makes your script faster.

2. Getting All Reviews (Handling Infinite Scroll)

Rvest can only parse static HTML, so it won't pick up content loaded via JavaScript when you scroll. You have two solid options here:

Option 1: Use Steam's Hidden API (Faster & More Reliable)

Steam loads reviews via an AJAX API endpoint, which we can call directly with httr to fetch batches of reviews. This avoids needing to simulate a browser.

Here's a script that loops through all pages of reviews:

library(rvest)
library(httr)
library(stringr)
library(dplyr)

app_id <- "215080"
base_url <- "https://steamcommunity.com/app/%s/homecontent/"

# Initialize storage for all reviews
all_reviews <- list()
offset <- 0
batch_size <- 10 # Steam returns 10 reviews per request by default

while(TRUE) {
  # Build the API request with pagination parameters
  url <- sprintf(base_url, app_id)
  params <- list(
    userreviewsoffset = offset,
    numperpage = batch_size,
    filter = "all",
    sort = "mostrecent", # Change to "mosthelpful" if needed
    language = "english", # Adjust to your target language
    reviewType = "all"
  )
  
  # Send the request
  response <- GET(url, query = params)
  if (http_status(response)$category != "Success") break
  
  # Parse the returned HTML
  page <- read_html(content(response, "text"))
  
  # Extract data for this batch
  date <- page %>% html_nodes("div.date_posted") %>% html_text() %>% str_trim()
  raw_review <- page %>% html_nodes("div.apphub_CardTextContent") %>% html_text()
  review <- map2_chr(raw_review, date, ~str_remove_all(.x, .y) %>% str_trim())
  opinion <- page %>% html_nodes("div.title") %>% html_text() %>% str_trim()
  hoursplayed <- page %>% html_nodes("div.hours") %>% html_text() %>% str_trim()
  helpful <- page %>% html_nodes("div.found_helpful") %>% html_text() %>% str_trim()
  
  # Stop if we hit an empty page
  if (length(date) == 0) break
  
  # Add this batch to our list
  all_reviews[[offset/batch_size + 1]] <- data.frame(
    Posted = date,
    Review = review,
    Opinion = opinion,
    Hours_Played = hoursplayed,
    Helpful_Votes = helpful
  )
  
  # Move to the next batch
  offset <- offset + batch_size
  Sys.sleep(2) # Add a small delay to avoid getting blocked by Steam
}

# Combine all batches into one data frame
final_reviews <- bind_rows(all_reviews)

Option 2: Simulate Browser Scrolling with RSelenium

If you prefer to mimic human scrolling (e.g., if the API parameters change), you can use RSelenium to control a browser:

library(RSelenium)
library(rvest)
library(stringr)

# Start a Chrome driver (you'll need ChromeDriver installed first)
driver <- rsDriver(browser = "chrome", port = 4567L)
remote_driver <- driver[["client"]]

# Navigate to the reviews page
remote_driver$navigate("https://steamcommunity.com/app/215080/reviews/")

# Scroll to load all reviews
last_height <- remote_driver$executeScript("return document.body.scrollHeight;")
while(TRUE) {
  remote_driver$executeScript("window.scrollTo(0, document.body.scrollHeight);")
  Sys.sleep(3) # Wait for new content to load
  new_height <- remote_driver$executeScript("return document.body.scrollHeight;")
  if (new_height == last_height) break # Stop when no new content loads
  last_height <- new_height
}

# Extract the full page HTML
page <- read_html(remote_driver$getPageSource()[[1]])

# Now extract your data using the same method as before
date <- page %>% html_nodes("div.date_posted") %>% html_text() %>% str_trim()
raw_review <- page %>% html_nodes("div.apphub_CardTextContent") %>% html_text()
review <- map2_chr(raw_review, date, ~str_remove_all(.x, .y) %>% str_trim())
# ... extract other fields here

# Clean up
remote_driver$close()
driver$server$stop()

This method is slower but works if the API approach isn't feasible.

内容的提问来源于stack exchange,提问作者KudoJr

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:44:09