如何轻松将特殊字符替换为rvest兼容的UTF-8(十六进制)格式?
rvest::read_html()'s Issue with URLs Containing Special Characters You don’t have to wait for errors and manually look up each special character’s percent-encoded value—R has built-in tools to automate this process entirely, converting characters like é to %C3%A9 and â to %C3%A2 in one go.
Recommended Method: Use httr::url_encode()
The httr package’s url_encode() function is designed to handle proper RFC-compliant URL encoding. It will encode all non-ASCII characters while leaving URL-specific separators (like /, ?, &) intact—exactly what read_html() expects.
Here’s how to apply it to your list of URLs with dplyr and purrr:
# Install and load required packages if you haven't already install.packages("httr") library(httr) library(dplyr) library(purrr) # Encode every URL in your list encoded_urls <- list_of_urls %>% map_chr(url_encode) # Now safely pass the encoded URLs to read_html() scraped_pages <- encoded_urls %>% map(read_html)
Alternative: Base R's utils::URLencode()
If you prefer not to install an extra package, you can use base R’s URLencode() function. Just make sure to set reserved = TRUE to keep URL structure characters unencoded:
library(dplyr) library(purrr) encoded_urls <- list_of_urls %>% map_chr(function(url) URLencode(url, reserved = TRUE))
Bonus: Encoding Specific Parts of a URL
If only a segment of your URL has special characters (like a path component), you can encode just that part before building the full URL:
base_url <- "https://example.com/menu/" menu_item <- "café crème" encoded_item <- url_encode(menu_item) full_encoded_url <- paste0(base_url, encoded_item) # Result: "https://example.com/menu/caf%C3%A9%20cr%C3%A8me"
This approach eliminates the need for manual gsub() fixes and ensures all your URLs are compatible with rvest right from the start.
内容的提问来源于stack exchange,提问作者stevec

