对总资产变量取log产生NaN的原因及R批量处理代码求助
Hey there! Let's tackle your two questions step by step—first figuring out that confusing NaN warning, then streamlining your code to batch create those log variables.
log() It’s totally reasonable to be confused when you’re explicitly filtering for values >0 but still getting NaNs. Here are the most likely culprits and how to diagnose them:
Hidden non-positive values (thanks to floating-point quirks)
Sometimes numbers that look positive are actually tiny negative values (like-1e-16) due to floating-point calculation errors, or they might beInf/-Infwhich also return NaN when logged. Run these checks to spot them:# Get a quick overview of your variable's distribution summary(HRSdata2$h1atotb) # Find rows where h1atotb > 0 but log() returns NaN bad_rows <- which(HRSdata2$h1atotb > 0 & is.na(log(HRSdata2$h1atotb))) # Inspect the actual values in those problematic rows HRSdata2$h1atotb[bad_rows] # Double-check that all values marked as >0 are truly positive all(HRSdata2$h1atotb[HRSdata2$h1atotb > 0] > 0)Check your data type
Ifh1atotbis stored as a character vector instead of numeric,log()will throw errors—but sometimes mixed-type columns (with hidden character values) can cause unexpected NaNs. Verify the type and look for anomalies:class(HRSdata2$h1atotb) which(!is.numeric(HRSdata2$h1atotb))Tweak your
ifelse()for robustness
While the root cause is the problematic values, you can make your code more resilient by explicitly handling NA values too:HRSdata2$h1logass <- ifelse(HRSdata2$h1atotb > 0 & !is.na(HRSdata2$h1atotb), log(HRSdata2$h1atotb), 0)
No need to repeat code 10 times! Here are two clean methods to automate this:
Base R: For Loop (simple and easy to follow)
This uses basic R syntax without extra packages:
# Define your original column names and corresponding new log column names orig_cols <- paste0("h", 1:10, "atotb") new_cols <- paste0("h", 1:10, "logass") # Loop through each pair of columns for (i in seq_along(orig_cols)) { HRSdata2[[new_cols[i]]] <- ifelse(HRSdata2[[orig_cols[i]]] > 0 & !is.na(HRSdata2[[orig_cols[i]]]), log(HRSdata2[[orig_cols[i]]]), 0) }
Using [[ instead of $ ensures the loop correctly references column names stored as strings.
Tidyverse: dplyr::across() (cleaner for data workflows)
If you use the tidyverse, this method is concise and readable:
library(dplyr) library(stringr) # For string replacement HRSdata2 <- HRSdata2 %>% mutate( # Target all columns from h1atotb to h10atotb across(paste0("h", 1:10, "atotb"), # Apply your log logic to each column ~ifelse(.x > 0 & !is.na(.x), log(.x), 0), # Automatically name new columns by replacing "atotb" with "logass" .names = "{str_replace(.col, 'atotb', 'logass')}") )
内容的提问来源于stack exchange,提问作者Emil Krabbe

