如何用R语言rvest包提取指定单个文本?排查并修正代码错误
The problem with your code is that the id you're targeting (freeText7754653807125407088) is dynamic—Goodreads generates a unique, random suffix for this ID every time the page loads. That means the exact ID you copied won’t exist when you run your code again, or for other users accessing the page.
Corrected Code
Instead of relying on the dynamic ID, use a more robust selector that targets the review text based on static class names or structural hierarchy. Here are two reliable approaches:
Approach 1: Target by Static Class Hierarchy
Use the static reviewText class (which wraps each review) and the freeText class (which contains the actual review text):
library(rvest) url <- "https://www.goodreads.com/book/show/3.Harry_Potter_and_the_Sorcerer_s_Stone?from_search=true&from_srp=true&qid=6Kzw6iKKur&rank=2#other_reviews" page <- read_html(url) # Extract all review texts review_texts <- page %>% html_nodes('.reviewText .freeText') %>% html_text(trim = TRUE) # Get the first review (adjust the index for others) first_review <- review_texts[1] cat(first_review)
Approach 2: Target by ID Pattern with XPath
If you want to specifically target spans whose IDs start with freeText (matching the dynamic ID pattern), use an XPath selector:
review_texts <- page %>% html_nodes(xpath = '//div[@class="reviewText"]//span[starts-with(@id, "freeText")]') %>% html_text(trim = TRUE)
Key Notes
trim = TRUEremoves extra whitespace from the extracted text, making it cleaner.- If you need a specific review (not just the first), filter by position (e.g.,
review_texts[3]for the third review) or add additional selectors (like targeting a specific user’s review by their username). - Most modern websites use dynamic IDs to prevent easy scraping, so always prioritize static class names or structural selectors over specific IDs when web scraping.
内容的提问来源于stack exchange,提问作者Christina

