线性模型双变量对数变换出现‘NaNs produced’错误的解决咨询
Hey there! That NaNs produced error is telling you exactly what's going wrong: your log() calls are trying to calculate the logarithm of 0 or negative numbers, which isn't mathematically defined (logarithms only exist for positive values). Let's walk through how to diagnose and fix this step by step.
Step 1: Identify the Problem Data
First, let's confirm which values are causing the issue by counting how many non-positive entries are in your variables:
# Count non-positive values in GINF sum(reg.conv$GINF <= 0, na.rm = TRUE) # Count non-positive values in INFt2014 sum(reg.conv$INFt2014 <= 0, na.rm = TRUE)
The na.rm = TRUE argument ignores any existing NA values so you get an accurate count of problematic 0/negative entries.
Step 2: Choose a Fix Based on Your Data
Now that you know where the issue is, pick one of these solutions depending on whether the non-positive values are valid data or errors:
Option 1: Remove Rows with Non-Positive Values (If They're Errors)
If those 0/negative entries are typos or invalid observations, simply filter them out before fitting your model:
# Keep only rows where both variables are positive reg.conv_clean <- reg.conv[reg.conv$GINF > 0 & reg.conv$INFt2014 > 0, ] # Refit your model and plot with the cleaned data model.conv <- lm(data = reg.conv_clean, log(GINF) ~ log(INFt2014)) x <- log(reg.conv_clean$GINF) y <- log(reg.conv_clean$INFt2014) plot(x, y, pch = 19, cex = 1.5, ylim = c(0,10), xlab = "LN_INFt2014", ylab = "LN_GINF", main = "Scatter plot between regional inflation in 2016 and the Growth of regional inflation from 2014 to 2016")
Just make sure to note in your analysis that you removed these rows, and check that your sample size doesn't drop too much to maintain statistical power.
Option 2: Add a Small Constant (If Non-Positive Values Are Valid)
If those non-positive values are legitimate (e.g., zero growth), you can shift all values by a tiny positive number to make them log-transformable. Choose an epsilon (eps) that's small enough not to distort your data but large enough to turn non-positives into positives:
# Pick a small offset (adjust based on your data scale) eps <- 0.01 # Use shifted values for log transformation model.conv <- lm(data = reg.conv, log(GINF + eps) ~ log(INFt2014 + eps)) x <- log(reg.conv$GINF + eps) y <- log(reg.conv$INFt2014 + eps) plot(x, y, pch = 19, cex = 1.5, ylim = c(0,10), xlab = "LN_INFt2014 (shifted)", ylab = "LN_GINF (shifted)", main = "Scatter plot between regional inflation in 2016 and the Growth of regional inflation from 2014 to 2016")
Be sure to report the offset value in your results so others can replicate your work.
Option 3: Use a More Flexible Transformation (Box-Cox)
If you want a statistically optimal transformation that handles non-positives automatically, try the Box-Cox transformation. It finds the best power transformation for your data (and reduces to log transformation when appropriate):
# Load the MASS package (comes with base R but needs to be loaded) library(MASS) # Find optimal lambda for GINF bc_ginf <- boxcox(GINF ~ INFt2014, data = reg.conv, plotit = FALSE) lambda_ginf <- bc_ginf$x[which.max(bc_ginf$y)] # Find optimal lambda for INFt2014 bc_inf <- boxcox(INFt2014 ~ GINF, data = reg.conv, plotit = FALSE) lambda_inf <- bc_inf$x[which.max(bc_inf$y)] # Fit model with Box-Cox transformed variables model.conv <- lm(data = reg.conv, (GINF^lambda_ginf - 1)/lambda_ginf ~ (INFt2014^lambda_inf - 1)/lambda_inf) # Plot transformed values x <- (reg.conv$INFt2014^lambda_inf - 1)/lambda_inf y <- (reg.conv$GINF^lambda_ginf - 1)/lambda_ginf plot(x, y, pch = 19, cex = 1.5, ylim = c(0,10), xlab = "Box-Cox transformed INFt2014", ylab = "Box-Cox transformed GINF", main = "Scatter plot between transformed regional inflation in 2016 and transformed growth of regional inflation (2014-2016)")
When lambda is 0, this is exactly the log transformation—so if the optimal lambda is close to 0, you can still interpret it similarly to a log model.
内容的提问来源于stack exchange,提问作者D. Smel

