You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言中数值型变量相关性分析返回NaN的问题排查

问题描述

处理多变量数据框时,此前生成相关矩阵均正常。新增由两个变量比值得到的变量ratiohighlow后,计算其与TotIncidents的相关性时,R返回NaN。所有相关变量均为数值型,执行is.nan(ratiohighlow)返回全FALSE,尝试限制该变量小数位数也无效果。

相关代码:

S1901Income_2022_crimetest <- S1901Income2022 |>
  select(HouseholdsTotal, Pct0to10k, Pct10kto14999, Pct15kto24999...8, Pct25kto34999, Pct35kto49999, Pct50kto74999, Pct75kto99999, Pct100kto149999, Pct150kto199999, Pct200korMore...22, HouseholdMedianIncome)

ratiohighlow <- (S1901Income_2022_crimetest$Pct200korMore...22 / S1901Income_2022_crimetest$Pct0to10k)
ratiohighlow <- round(ratiohighlow, digits = 8)

Crime2022_S1901Income <- Crime2022 |>
  select(-FIPS, -year)

IncomeandCrime2022 <- cbind(Crime2022_S1901Income, S1901Income_2022_crimetest, ratiohighlow)

cor(IncomeandCrime2022$TotIncidents, ratiohighlow)

数据样本:

# A tibble: 10 × 12
   HouseholdsTotal Pct0to10k Pct10kto14999 Pct15kto24999...8 Pct25kto34999 Pct35kto49999 Pct50kto74999
             <dbl>     <dbl>         <dbl>             <dbl>         <dbl>         <dbl>         <dbl>
 1          265794       6.6           5                 8.6           8.4          11.9          16.7
 2         1665560       4.2           2.5               5.5           6.6          10.6          17.1
 3          151490       4.4           2.2               6             7.2          11.4          20.2
 4           74678       6.9           4.5               9.3           9.8          14.1          20.2
 5            6559       6.2           7.4              13.1           7.2           7.9          22.4
 6            7194       9             8                16             7.8          13.3          18.6
 7           19021       5.3           3.6              14.7          12.8          16.8          20.2
 8          104164       2.9           2.2               5.3           6            10.2          17.3
 9           15172       5.7           5                10.9          10.8          15.1          20.6
10            3587       8.8          10.2              15             8.7          13.6          17.9
# ℹ 5 more variables: Pct75kto99999 <dbl>, Pct100kto149999 <dbl>, Pct150kto199999 <dbl>, Pct200korMore...22 <dbl>,
#   HouseholdMedianIncome <dbl>

补充数据:

Pct100kto149999 <dbl>
Pct150kto199999 <dbl>
Pct200korMore...22 <dbl>
HouseholdMedianIncome <dbl>
ratiohighlow <dbl>
14.1    7.3 9.3 63595   1.4090909
18.6    9.5 11.3    80675   2.6904762
19.4    7.6 6.3 73313   1.4318182
13.3    4.8 4.1 56439   0.5942029
14.0    5.0 2.6 58695   0.4193548
12.0    2.4 2.2 44804   0.2444444
可能的原因与解决办法
  • 检查分母为0的情况:虽然is.nan(ratiohighlow)返回全FALSE,但如果Pct0to10k存在0值,计算比值会得到Inf(无穷大),cor函数遇到Inf会返回NaN。执行以下代码排查:

    sum(S1901Income_2022_crimetest$Pct0to10k == 0)
    

    若存在0值,可将对应行的ratiohighlow替换为NA,或直接删除这些行:

    ratiohighlow[S1901Income_2022_crimetest$Pct0to10k == 0] <- NA
    
  • 检查TotIncidents的异常值:TotIncidents中可能存在NA或Inf,导致相关性计算失败。执行以下代码排查:

    sum(is.na(IncomeandCrime2022$TotIncidents))
    sum(is.infinite(IncomeandCrime2022$TotIncidents))
    

    若存在异常值,需先处理(如填充或删除)。

  • 使用cor函数的use参数:添加use="complete.obs"参数,忽略包含缺失值或无穷大的行,直接计算有效样本的相关性:

    cor(IncomeandCrime2022$TotIncidents, ratiohighlow, use="complete.obs")
    
  • 确认数据合并的对齐性:用cbind合并数据框时,若两个数据框行数不一致,会自动循环补齐导致数据错位。检查行数是否匹配:

    nrow(Crime2022_S1901Income) == nrow(S1901Income_2022_crimetest)
    

    建议使用dplyr::bind_cols替代cbind,合并时会严格检查行数一致性:

    library(dplyr)
    IncomeandCrime2022 <- bind_cols(Crime2022_S1901Income, S1901Income_2022_crimetest, tibble(ratiohighlow))
    

内容的提问来源于stack exchange,提问作者Steven Morrison

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 19:29:55