R语言中数值型变量相关性分析返回NaN的问题排查
问题描述
处理多变量数据框时,此前生成相关矩阵均正常。新增由两个变量比值得到的变量ratiohighlow后,计算其与TotIncidents的相关性时,R返回NaN。所有相关变量均为数值型,执行is.nan(ratiohighlow)返回全FALSE,尝试限制该变量小数位数也无效果。
相关代码:
S1901Income_2022_crimetest <- S1901Income2022 |> select(HouseholdsTotal, Pct0to10k, Pct10kto14999, Pct15kto24999...8, Pct25kto34999, Pct35kto49999, Pct50kto74999, Pct75kto99999, Pct100kto149999, Pct150kto199999, Pct200korMore...22, HouseholdMedianIncome) ratiohighlow <- (S1901Income_2022_crimetest$Pct200korMore...22 / S1901Income_2022_crimetest$Pct0to10k) ratiohighlow <- round(ratiohighlow, digits = 8) Crime2022_S1901Income <- Crime2022 |> select(-FIPS, -year) IncomeandCrime2022 <- cbind(Crime2022_S1901Income, S1901Income_2022_crimetest, ratiohighlow) cor(IncomeandCrime2022$TotIncidents, ratiohighlow)
数据样本:
# A tibble: 10 × 12 HouseholdsTotal Pct0to10k Pct10kto14999 Pct15kto24999...8 Pct25kto34999 Pct35kto49999 Pct50kto74999 <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> 1 265794 6.6 5 8.6 8.4 11.9 16.7 2 1665560 4.2 2.5 5.5 6.6 10.6 17.1 3 151490 4.4 2.2 6 7.2 11.4 20.2 4 74678 6.9 4.5 9.3 9.8 14.1 20.2 5 6559 6.2 7.4 13.1 7.2 7.9 22.4 6 7194 9 8 16 7.8 13.3 18.6 7 19021 5.3 3.6 14.7 12.8 16.8 20.2 8 104164 2.9 2.2 5.3 6 10.2 17.3 9 15172 5.7 5 10.9 10.8 15.1 20.6 10 3587 8.8 10.2 15 8.7 13.6 17.9 # ℹ 5 more variables: Pct75kto99999 <dbl>, Pct100kto149999 <dbl>, Pct150kto199999 <dbl>, Pct200korMore...22 <dbl>, # HouseholdMedianIncome <dbl>
补充数据:
Pct100kto149999 <dbl> Pct150kto199999 <dbl> Pct200korMore...22 <dbl> HouseholdMedianIncome <dbl> ratiohighlow <dbl> 14.1 7.3 9.3 63595 1.4090909 18.6 9.5 11.3 80675 2.6904762 19.4 7.6 6.3 73313 1.4318182 13.3 4.8 4.1 56439 0.5942029 14.0 5.0 2.6 58695 0.4193548 12.0 2.4 2.2 44804 0.2444444
可能的原因与解决办法
检查分母为0的情况:虽然
is.nan(ratiohighlow)返回全FALSE,但如果Pct0to10k存在0值,计算比值会得到Inf(无穷大),cor函数遇到Inf会返回NaN。执行以下代码排查:sum(S1901Income_2022_crimetest$Pct0to10k == 0)若存在0值,可将对应行的
ratiohighlow替换为NA,或直接删除这些行:ratiohighlow[S1901Income_2022_crimetest$Pct0to10k == 0] <- NA检查
TotIncidents的异常值:TotIncidents中可能存在NA或Inf,导致相关性计算失败。执行以下代码排查:sum(is.na(IncomeandCrime2022$TotIncidents)) sum(is.infinite(IncomeandCrime2022$TotIncidents))若存在异常值,需先处理(如填充或删除)。
使用
cor函数的use参数:添加use="complete.obs"参数,忽略包含缺失值或无穷大的行,直接计算有效样本的相关性:cor(IncomeandCrime2022$TotIncidents, ratiohighlow, use="complete.obs")确认数据合并的对齐性:用
cbind合并数据框时,若两个数据框行数不一致,会自动循环补齐导致数据错位。检查行数是否匹配:nrow(Crime2022_S1901Income) == nrow(S1901Income_2022_crimetest)建议使用
dplyr::bind_cols替代cbind,合并时会严格检查行数一致性:library(dplyr) IncomeandCrime2022 <- bind_cols(Crime2022_S1901Income, S1901Income_2022_crimetest, tibble(ratiohighlow))
内容的提问来源于stack exchange,提问作者Steven Morrison
相关产品推荐
相关产品推荐

