使用R抓取HTML表格(含iea.org)为何得到字符串而非整数?
Hey there! I’ve run into this exact problem before when pulling energy data from the IEA site—let’s break down why this happens and how to fix it.
Why You’re Getting Strings Instead of Numbers
There are three main reasons this crops up:
- Non-numeric extra characters: IEA tables often add thousand separators (like
,), units (e.g.,MWh,USD), or hidden whitespace to number cells. Tools likervest::html_table()default to treating these as strings because they can’t automatically guess which extra content is safe to remove. - Nested HTML structure: Numbers are sometimes wrapped in
<span>or other styling tags.html_table()grabs the full text of the cell, but doesn’t dig into nested elements to extract just the numeric value. - Tool default behavior:
html_table()is built to preserve raw HTML text to avoid accidental misconversion (e.g., turning a numeric ID into a measurable value). It won’t convert types unless you explicitly tell it to.
Step-by-Step Fixes
Let’s use rvest (the go-to package for this task) with examples tailored to IEA’s tables.
1. Start with Basic Scraping & Check Types
First, confirm what you’re working with:
library(rvest) library(dplyr) # Replace with your target IEA page URL target_url <- "https://www.iea.org/your-target-page" page <- read_html(target_url) # Grab the first table on the page (adjust the index if you need a different one) raw_table <- page %>% html_table(fill = TRUE) %>% .[[1]] # Check column types—you'll see strings where you expect numbers str(raw_table)
2. Auto-Clean & Convert with readr::parse_number()
This is the easiest fix for most IEA tables, as it automatically strips commas, symbols, and whitespace:
library(readr) # Convert specific columns (e.g., columns 2 through 6) to numbers clean_table <- raw_table %>% mutate(across(2:6, parse_number)) # Or convert by column name (replace with your actual column names) clean_table <- raw_table %>% mutate(across(c("2020", "2021", "2022"), parse_number))
3. Handle Units or Custom Text
If cells include fixed units (like MWh), use stringr to strip them first:
library(stringr) clean_table <- raw_table %>% mutate( # Remove "MWh" and surrounding whitespace, then convert to numeric Renewable_Energy = str_remove(Renewable_Energy, "\\s*MWh$") %>% as.numeric(), # For currency values, remove "$" and commas Investment = str_remove(Investment, "^\\$") %>% str_remove_all(",") %>% as.numeric() )
4. Dig into Nested HTML (If Needed)
If html_table() isn’t grabbing clean numbers because of nested tags, target the specific elements holding the numbers directly:
# Grab all table rows table_rows <- page %>% html_elements("table tr") # Extract headers table_headers <- table_rows %>% first() %>% html_elements("th") %>% html_text(trim = TRUE) # Extract numeric values from nested span tags (common in IEA's styled tables) clean_rows <- table_rows %>% tail(-1) %>% # Skip the header row map(function(row) { row %>% html_elements("td span[data-value]") %>% # Target the tag with raw numbers html_attr("data-value") %>% # Grab the raw numeric attribute instead of displayed text as.numeric() }) # Convert to a clean data frame clean_table <- data.frame(do.call(rbind, clean_rows)) colnames(clean_table) <- table_headers
Final Tip
Always inspect the HTML source of the IEA table (right-click > "Inspect" in your browser) to see exactly how numbers are stored. That’ll tell you whether you need to strip characters, target nested tags, or use attribute values instead of displayed text.
内容的提问来源于stack exchange,提问作者RawisWar

