You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用R抓取HTML表格(含iea.org)为何得到字符串而非整数?

Fixing String-to-Number Issues When Scraping IEA Tables in R

Hey there! I’ve run into this exact problem before when pulling energy data from the IEA site—let’s break down why this happens and how to fix it.

Why You’re Getting Strings Instead of Numbers

There are three main reasons this crops up:

  • Non-numeric extra characters: IEA tables often add thousand separators (like ,), units (e.g., MWh, USD), or hidden whitespace to number cells. Tools like rvest::html_table() default to treating these as strings because they can’t automatically guess which extra content is safe to remove.
  • Nested HTML structure: Numbers are sometimes wrapped in <span> or other styling tags. html_table() grabs the full text of the cell, but doesn’t dig into nested elements to extract just the numeric value.
  • Tool default behavior: html_table() is built to preserve raw HTML text to avoid accidental misconversion (e.g., turning a numeric ID into a measurable value). It won’t convert types unless you explicitly tell it to.

Step-by-Step Fixes

Let’s use rvest (the go-to package for this task) with examples tailored to IEA’s tables.

1. Start with Basic Scraping & Check Types

First, confirm what you’re working with:

library(rvest)
library(dplyr)

# Replace with your target IEA page URL
target_url <- "https://www.iea.org/your-target-page"
page <- read_html(target_url)

# Grab the first table on the page (adjust the index if you need a different one)
raw_table <- page %>% html_table(fill = TRUE) %>% .[[1]]

# Check column types—you'll see strings where you expect numbers
str(raw_table)

2. Auto-Clean & Convert with readr::parse_number()

This is the easiest fix for most IEA tables, as it automatically strips commas, symbols, and whitespace:

library(readr)

# Convert specific columns (e.g., columns 2 through 6) to numbers
clean_table <- raw_table %>%
  mutate(across(2:6, parse_number))

# Or convert by column name (replace with your actual column names)
clean_table <- raw_table %>%
  mutate(across(c("2020", "2021", "2022"), parse_number))

3. Handle Units or Custom Text

If cells include fixed units (like MWh), use stringr to strip them first:

library(stringr)

clean_table <- raw_table %>%
  mutate(
    # Remove "MWh" and surrounding whitespace, then convert to numeric
    Renewable_Energy = str_remove(Renewable_Energy, "\\s*MWh$") %>% as.numeric(),
    # For currency values, remove "$" and commas
    Investment = str_remove(Investment, "^\\$") %>% str_remove_all(",") %>% as.numeric()
  )

4. Dig into Nested HTML (If Needed)

If html_table() isn’t grabbing clean numbers because of nested tags, target the specific elements holding the numbers directly:

# Grab all table rows
table_rows <- page %>% html_elements("table tr")

# Extract headers
table_headers <- table_rows %>% 
  first() %>% 
  html_elements("th") %>% 
  html_text(trim = TRUE)

# Extract numeric values from nested span tags (common in IEA's styled tables)
clean_rows <- table_rows %>% 
  tail(-1) %>% # Skip the header row
  map(function(row) {
    row %>% 
      html_elements("td span[data-value]") %>% # Target the tag with raw numbers
      html_attr("data-value") %>% # Grab the raw numeric attribute instead of displayed text
      as.numeric()
  })

# Convert to a clean data frame
clean_table <- data.frame(do.call(rbind, clean_rows))
colnames(clean_table) <- table_headers

Final Tip

Always inspect the HTML source of the IEA table (right-click > "Inspect" in your browser) to see exactly how numbers are stored. That’ll tell you whether you need to strip characters, target nested tags, or use attribute values instead of displayed text.

内容的提问来源于stack exchange,提问作者RawisWar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 08:51:36