You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用R语言rvest包提取指定单个文本?排查并修正代码错误

Fixing Dynamic ID Issue When Extracting Review Text from Goodreads with rvest

The problem with your code is that the id you're targeting (freeText7754653807125407088) is dynamic—Goodreads generates a unique, random suffix for this ID every time the page loads. That means the exact ID you copied won’t exist when you run your code again, or for other users accessing the page.

Corrected Code

Instead of relying on the dynamic ID, use a more robust selector that targets the review text based on static class names or structural hierarchy. Here are two reliable approaches:

Approach 1: Target by Static Class Hierarchy

Use the static reviewText class (which wraps each review) and the freeText class (which contains the actual review text):

library(rvest)

url <- "https://www.goodreads.com/book/show/3.Harry_Potter_and_the_Sorcerer_s_Stone?from_search=true&from_srp=true&qid=6Kzw6iKKur&rank=2#other_reviews"
page <- read_html(url)

# Extract all review texts
review_texts <- page %>% 
  html_nodes('.reviewText .freeText') %>% 
  html_text(trim = TRUE)

# Get the first review (adjust the index for others)
first_review <- review_texts[1]
cat(first_review)

Approach 2: Target by ID Pattern with XPath

If you want to specifically target spans whose IDs start with freeText (matching the dynamic ID pattern), use an XPath selector:

review_texts <- page %>% 
  html_nodes(xpath = '//div[@class="reviewText"]//span[starts-with(@id, "freeText")]') %>% 
  html_text(trim = TRUE)

Key Notes

  • trim = TRUE removes extra whitespace from the extracted text, making it cleaner.
  • If you need a specific review (not just the first), filter by position (e.g., review_texts[3] for the third review) or add additional selectors (like targeting a specific user’s review by their username).
  • Most modern websites use dynamic IDs to prevent easy scraping, so always prioritize static class names or structural selectors over specific IDs when web scraping.

内容的提问来源于stack exchange,提问作者Christina

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 07:39:43