You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言中为百分比列添加%符号并去除NaN值的方法

解决R中百分比列添加%符号同时替换NaN为0的问题

问题描述

使用以下代码生成数据框tab9时,希望为mutate生成的5个百分比列添加%符号,但直接用paste0会把NaN转为'NaN%'字符串,导致后续tab9[is.na(tab9)] <- 0无法替换这些值:

tab9 <- table(tab9data %>%
  select(currentGrade, percentDiffRange) %>%
  group_by(currentGrade)) %>%
  as.data.frame() %>%
  pivot_wider(names_from = 'percentDiffRange', values_from = 'Freq') %>%
  mutate(perc0 = round((tab9$`0%`/ rowSums(tab9[ ,2:6])) * 100,1),
         perc0.1_1.9 = round((tab9$`0.1% - 1.9%`/ rowSums(tab9[ ,2:6])) * 100,1),
         perc2_3.9 = round((tab9$`2% - 3.9%`/ rowSums(tab9[ ,2:6])) * 100,1),
         perc4_5.9 = round((tab9$`4% - 5.9%`/ rowSums(tab9[ ,2:6])) * 100,1),
         perc6 = round((tab9$`6% or more`/ rowSums(tab9[ ,2:6])) * 100,1))

tab9[is.na(tab9)] <- 0

解决方案

核心思路是先处理数值型的NaN/NA,再转换为带%的字符串,同时修正原代码中在mutate里引用未生成对象的错误:

方法一:先替换NaN为0,再添加%符号

tab9 <- tab9data %>%
  select(currentGrade, percentDiffRange) %>%
  table() %>%
  as.data.frame() %>%
  pivot_wider(names_from = 'percentDiffRange', values_from = 'Freq') %>%
  # 计算百分比(数值型,暂不添加%)
  mutate(
    perc0 = round((`0%`/ rowSums(select(., 2:6))) * 100, 1),
    perc0.1_1.9 = round((`0.1% - 1.9%`/ rowSums(select(., 2:6))) * 100, 1),
    perc2_3.9 = round((`2% - 3.9%`/ rowSums(select(., 2:6))) * 100, 1),
    perc4_5.9 = round((`4% - 5.9%`/ rowSums(select(., 2:6))) * 100, 1),
    perc6 = round((`6% or more`/ rowSums(select(., 2:6))) * 100, 1)
  ) %>%
  # 将数值型的NaN/NA替换为0
  mutate(across(c(perc0, perc0.1_1.9, perc2_3.9, perc4_5.9, perc6), ~ifelse(is.na(.), 0, .))) %>%
  # 为百分比列添加%符号,转为字符串型
  mutate(across(c(perc0, perc0.1_1.9, perc2_3.9, perc4_5.9, perc6), ~paste0(., "%")))

方法二:生成字符串时直接处理NaN

如果想一步完成计算和格式转换,可以在paste0前用ifelse判断分母为0的情况(这是NaN产生的常见原因),直接返回"0%":

tab9 <- tab9data %>%
  select(currentGrade, percentDiffRange) %>%
  table() %>%
  as.data.frame() %>%
  pivot_wider(names_from = 'percentDiffRange', values_from = 'Freq') %>%
  mutate(
    perc0 = paste0(round(ifelse(rowSums(select(., 2:6)) == 0, 0, (`0%`/ rowSums(select(., 2:6))) * 100), 1), "%"),
    perc0.1_1.9 = paste0(round(ifelse(rowSums(select(., 2:6)) == 0, 0, (`0.1% - 1.9%`/ rowSums(select(., 2:6))) * 100), 1), "%"),
    perc2_3.9 = paste0(round(ifelse(rowSums(select(., 2:6)) == 0, 0, (`2% - 3.9%`/ rowSums(select(., 2:6))) * 100), 1), "%"),
    perc4_5.9 = paste0(round(ifelse(rowSums(select(., 2:6)) == 0, 0, (`4% - 5.9%`/ rowSums(select(., 2:6))) * 100), 1), "%"),
    perc6 = paste0(round(ifelse(rowSums(select(., 2:6)) == 0, 0, (`6% or more`/ rowSums(select(., 2:6))) * 100), 1), "%")
  )

# 若仍有遗漏的NA字符串,可补充替换
tab9[is.na(tab9)] <- "0%"

关键说明

  • 原代码中mutate里用tab9$引用列是错误的,因为此时tab9还未生成,改用select(., 列范围)或直接列名即可。
  • 直接用paste0会把数值型NaN转为字符串"NaN%",而is.na()只能识别数值型的NA/NaN,无法识别字符串,所以必须先处理数值型的缺失值再转字符串。

内容的提问来源于stack exchange,提问作者atm1984

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 19:15:38