R语言中全角数字字符无法转为数值的解决求助
全角数字/符号转数值类型问题解决
问题背景
从中文PDF提取的数据中,数字和负号为全角格式(如-122、29458),字符类型列使用as.numeric()或parse_number()转换时全部返回NA。
提取数据的核心代码:
library(tidyverse) library(pdftools) file <- tempfile() url <- paste0("http://images.mofcom.gov.cn/fec/202211/20221118091910924.pdf") download.file(url, file, headers = c("User-Agent" = "My Custom User Agent")) pdf_data <- pdf_text(file) replace_spaces_and_commas <- function(x) { str_replace_all(x, "[ ,]", "") } pdf <- pdf_data[53:71] tab_pdf <- str_split(pdf, "\n") for (i in 1:19) { assign(paste0("tab_pdf_", i), tab_pdf[[i]]) } the_names <- c("country", "year_2013", "year_2014", "year_2015", "year_2016", "year_2017", "year_2018", "year_2019", "year_2020", "year_2021") pdf_clean1 <- tab_pdf_1[14:60] %>% str_trim %>% str_replace_all(",", "") %>% str_split("\\s{2,}", simplify = TRUE) %>% data.frame(stringsAsFactors = FALSE) %>% setNames(the_names) %>% mutate_all(.funs = replace_spaces_and_commas) %>% filter(country != "")
测试样本:
c("-122", "29458", "-74", "16357", "2", "-534")
解决方法
方法1:用stringi批量转换全角到半角
stringi包的stri_trans_general()支持直接将全角(宽字符)转换为半角(窄字符),无需逐个替换:
library(stringi) # 转换单个年份列 pdf_clean1$year_2013 <- as.numeric(stri_trans_general(pdf_clean1$year_2013, "Wide-Narrow")) # 批量转换所有年份列 year_cols <- startsWith(names(pdf_clean1), "year_") pdf_clean1[year_cols] <- lapply(pdf_clean1[year_cols], function(x) { as.numeric(stri_trans_general(x, "Wide-Narrow")) })
方法2:手动替换全角字符(无需额外包)
如果不想加载新包,可手动映射全角字符到半角:
# 定义转换函数 full_to_half <- function(x) { x <- str_replace_all(x, "-", "-") x <- str_replace_all(x, c( "0"="0", "1"="1", "2"="2", "3"="3", "4"="4", "5"="5", "6"="6", "7"="7", "8"="8", "9"="9" )) as.numeric(x) } # 应用到单列 pdf_clean1$year_2013 <- full_to_half(pdf_clean1$year_2013) # 批量处理年份列 year_cols <- startsWith(names(pdf_clean1), "year_") pdf_clean1[year_cols] <- lapply(pdf_clean1[year_cols], full_to_half)
方法3:整合到现有清洗流程
修改现有的replace_spaces_and_commas函数,在清洗阶段完成转换:
# 基于stringi的版本 replace_spaces_and_commas <- function(x) { x <- str_replace_all(x, "[ ,]", "") stri_trans_general(x, "Wide-Narrow") } # 或者手动替换版本 replace_spaces_and_commas <- function(x) { x <- str_replace_all(x, "[ ,]", "") x <- str_replace_all(x, "-", "-") str_replace_all(x, c( "0"="0", "1"="1", "2"="2", "3"="3", "4"="4", "5"="5", "6"="6", "7"="7", "8"="8", "9"="9" )) } # 之后执行原清洗代码,最后转数值 pdf_clean1[year_cols] <- lapply(pdf_clean1[year_cols], as.numeric)
验证结果
对测试样本执行转换:
test_vec <- c("-122", "29458", "-74", "16357", "2", "-534") as.numeric(stri_trans_general(test_vec, "Wide-Narrow")) # 输出:-122 29458 -74 16357 2 -534
内容的提问来源于stack exchange,提问作者folder
相关产品推荐
相关产品推荐

