R quantreg:分类调查数据的奇异性问题及求解问询
quantreg for Probability Heuristic Replication Let’s break down the problems you’re encountering when replicating Blass (2008)’s probability heuristic method with linear median regression, and walk through actionable fixes:
First Issue: Singular Design Matrix (Dummy Trap)
Your initial fix of manually generating dummy variables (instead of relying on factor() which includes all levels plus an intercept) was correct—this is the classic dummy trap where including all factor levels alongside an intercept creates perfect multicollinearity. But moving to manual dummies introduced new warnings/errors, which we’ll address next.
Second Issue: "Solution may be nonunique" & Backsolve Singular Matrix Errors
These are common in quantile regression, often tied to data structure, model specification, or the solver method. Here’s how to diagnose and fix:
1. Check for Residual Multicollinearity
Even after manual dummy creation, your factor(NoneOpt) variable might be linearly dependent with other terms. For example, if NoneOpt = 1 always corresponds to all other dummy variables being 0, this creates a redundant relationship with the intercept.
Run this to detect aliased terms:
alias(lm(LogProb ~ Delivery.costs_1 + Delivery.costs_2 + Minimum.order.quantity_1 + Minimum.order.quantity_2 + Delivery.window_1 + Delivery.window_2 + factor(NoneOpt), data=PE_dichomitzed))
- If an alias exists, remove the redundant variable. For example, if
NoneOptis redundant with the intercept, either dropfactor(NoneOpt)or remove the intercept with-1in your formula:LM.factor <- rq(LogProb ~ -1 + Delivery.costs_1 + Delivery.costs_2 + Minimum.order.quantity_1 + Minimum.order.quantity_2 + Delivery.window_1 + Delivery.window_2 + factor(NoneOpt), data=PE_dichomitzed, tau=0.5)
2. Switch Quantile Regression Solver Methods
The default method="br" (Barrodale-Roberts) solver can struggle with non-unique solutions. Try more robust solvers:
- Frisch-Newton (
method="fn"): Works better for larger datasets and avoids some non-uniqueness issues:LM.factor <- rq(LogProb ~ Delivery.costs_1 + Delivery.costs_2 + Minimum.order.quantity_1 + Minimum.order.quantity_2 + Delivery.window_1 + Delivery.window_2 + factor(NoneOpt), data=PE_dichomitzed, tau=0.5, method="fn") - Pairwise Frisch-Newton (
method="pfn"): Useful if you have many parameters relative to sample size.
3. Validate Data Quality & Transformations
Double-check your LogProb calculation:
- Ensure
Response[1]never equals 0 (this would createInf/NaNvalues inLogProb). Filter out any groups whereResponse[1] == 0before computing the transformation. - Check for duplicate
LogProbvalues—large numbers of duplicates can lead to non-unique quantile regression solutions. If duplicates are unavoidable, a tiny amount of jitter (e.g.,LogProb = LogProb + rnorm(nrow(PE_raw), 0, 1e-6)) can help, but use this cautiously as it alters your data.
4. Adjust Model Complexity
If your sample size is small relative to the number of parameters (7 dummies + intercept = 8 parameters), reduce model complexity:
- Merge rare factor levels (e.g., if
Delivery.costs_2has very few observations, combine it withDelivery.costs_1). - Drop variables with minimal predictive power (use exploratory analysis like boxplots of
LogProbacross each dummy level to assess impact).
5. Fix Summary() Backsolve Errors
When calling summary(), the default variance estimation relies on a Hessian matrix that can be singular. Use bootstrap-based standard errors instead to bypass this:
summary(LM.factor, se="boot", R=1000)
The R parameter sets the number of bootstrap replicates—adjust based on computational resources.
内容的提问来源于stack exchange,提问作者C. Schrijver

