You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用R的XML和RCurl包从HTML页面(含指定页面)提取表格为data.frame

Extract HTML Tables to Data Frame in R with XML & RCurl (Forbes Powerful Brands Example)

Hey there! Let’s break down exactly how to pull HTML tables into a data.frame using R’s XML and RCurl packages, with a step-by-step example for the Forbes Powerful Brands list you specified.

Step 1: Install & Load Required Packages

First off, make sure you have the necessary packages installed. If not, run this:

# Install packages (only need to do this once)
install.packages(c("XML", "RCurl"))

# Load the libraries for use
library(XML)
library(RCurl)

Step 2: Fetch the HTML Page Content

Next, we’ll use RCurl to grab the raw HTML from the Forbes URL. Sometimes SSL certificate checks can cause issues, so we’ll add ssl.verifypeer = FALSE to bypass that:

# Define the target URL
url <- "https://www.forbes.com/powerful-brands/list/#tab:rank.html"

# Fetch the HTML content
html_raw <- getURL(url, ssl.verifypeer = FALSE)

Step 3: Parse the HTML

Now we’ll use the XML package to parse the raw HTML into a structure R can work with:

# Parse the HTML content
parsed_html <- htmlParse(html_raw)

Step 4: Extract Tables from the Parsed HTML

The readHTMLTable() function will pull all tables from the parsed HTML into a list. We can check how many tables were extracted first:

# Extract all tables into a list
all_tables <- readHTMLTable(parsed_html)

# See how many tables we have
length(all_tables)

Step 5: Target the Forbes Brands Table & Convert to Data Frame

For the Forbes Powerful Brands page, the main ranking table is typically one of the first entries in the list. Let’s pull it out and convert it to a data.frame:

# Grab the target table (adjust the index if needed—start with 1 if unsure)
brands_df <- as.data.frame(all_tables[[1]])

# Check the first few rows to confirm we got the right table
head(brands_df)

Step 6: Clean Up the Data (Optional)

You might notice some messy columns or default headers. Here’s a quick way to tidy things up:

# Remove any completely empty columns
brands_df <- brands_df[, colSums(is.na(brands_df)) != nrow(brands_df)]

# Rename columns to something more intuitive (match the Forbes table headers)
colnames(brands_df) <- c("Rank", "Brand", "Brand Value (BUSD)", "YoY Change", "Country", "Sector")

# Check the cleaned data
head(brands_df)

A Quick Note

Web pages often change their structure over time. If the table index doesn’t work, use str(all_tables) to inspect the list and find the table with the columns you need.

内容的提问来源于stack exchange,提问作者Roy Sudip

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:43:12