使用RCurl抓取网页后,如何提取特定class的a标签href值
Hey there! Let's break down how to pull those specific links from your scraped HTML. You've already got the raw webpage content using RCurl::getURL(), now we just need to parse it and target the right elements.
Step 1: Use a HTML parsing library
The easiest way to handle this in R is with the rvest package—it's built for scraping and works seamlessly with raw HTML strings. First, make sure you have it installed and loaded:
# Install if you haven't already install.packages("rvest") library(rvest) library(RCurl) # Keep your existing library loaded
Step 2: Parse the raw HTML and extract links
Take the string you got from getURL(), convert it into an HTML document, then use CSS selectors to grab the links you want:
# Your existing code to fetch the page webpage <- getURL("https://somewebpage.com") # Convert the raw HTML string into a parseable document html_doc <- read_html(webpage) # Target all <a> tags with class "answer_permalink" and extract their href attributes target_hrefs <- html_doc |> html_elements(".answer_permalink") |> # CSS selector for class html_attr("href") # View the results target_hrefs
Alternative: Using the XML package
If you prefer using the older XML package instead, here's how you'd do it with XPath:
library(XML) # Parse the HTML doc <- htmlParse(webpage) # Use XPath to select <a> tags where class equals "answer_permalink" target_hrefs <- xpathSApply(doc, "//a[@class='answer_permalink']", xmlGetAttr, "href")
Quick note
If the webpage loads content dynamically (like with JavaScript), RCurl might not capture the full HTML you see in your browser. In that case, you'd need a tool like RSelenium or rvest with V8 to render the JavaScript first. But based on the HTML snippet you shared, static scraping should work just fine here.
内容的提问来源于stack exchange,提问作者Ronak Shah

