You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用RCurl抓取网页后,如何提取特定class的a标签href值

Hey there! Let's break down how to pull those specific links from your scraped HTML. You've already got the raw webpage content using RCurl::getURL(), now we just need to parse it and target the right elements.

Step 1: Use a HTML parsing library

The easiest way to handle this in R is with the rvest package—it's built for scraping and works seamlessly with raw HTML strings. First, make sure you have it installed and loaded:

# Install if you haven't already
install.packages("rvest")
library(rvest)
library(RCurl) # Keep your existing library loaded

Take the string you got from getURL(), convert it into an HTML document, then use CSS selectors to grab the links you want:

# Your existing code to fetch the page
webpage <- getURL("https://somewebpage.com")

# Convert the raw HTML string into a parseable document
html_doc <- read_html(webpage)

# Target all <a> tags with class "answer_permalink" and extract their href attributes
target_hrefs <- html_doc |>
  html_elements(".answer_permalink") |> # CSS selector for class
  html_attr("href")

# View the results
target_hrefs

Alternative: Using the XML package

If you prefer using the older XML package instead, here's how you'd do it with XPath:

library(XML)

# Parse the HTML
doc <- htmlParse(webpage)

# Use XPath to select <a> tags where class equals "answer_permalink"
target_hrefs <- xpathSApply(doc, "//a[@class='answer_permalink']", xmlGetAttr, "href")

Quick note

If the webpage loads content dynamically (like with JavaScript), RCurl might not capture the full HTML you see in your browser. In that case, you'd need a tool like RSelenium or rvest with V8 to render the JavaScript first. But based on the HTML snippet you shared, static scraping should work just fine here.

内容的提问来源于stack exchange,提问作者Ronak Shah

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 03:55:42