基于非平衡面板数据的Nested Logit模型构建及策略排序技术求助(R语言)
Hey there! Let's work through your problem step by step, starting with your question about discretizing the dependent variable, then moving to building the right model for your strategy ranking goal.
First: Is Discretizing DiffPerformance Appropriate?
You converted your continuous performance measure to a binary DiffPerformanceHL (H/L), and you're right that this loses information about the magnitude of performance differences. That said:
- If your core goal is to rank strategies by their likelihood of delivering above-market performance, this binary variable works perfectly with logit-style models.
- If you also care about how much a strategy improves performance (not just whether it does), you might want to complement this with a panel regression (e.g., fixed-effects models using the
plmpackage) on the continuousDiffPerformancevariable. But since you asked about logit models, let's focus on the binary case first.
Clarifying Model Choice: Nested Logit vs. Hierarchical Logit
From your description, you want to rank strategies (nested in groups a/b) based on their impact on performance. Here's the key distinction:
- Nested Logit is typically used for discrete choice problems (e.g., "why do firms choose strategy X over Y?"). If your goal was to model how performance affects strategy choice, this would be the right fit.
- Hierarchical (Multilevel) Logit is better for your goal: it accounts for the nested structure of strategies (group a contains strategies 1-3, group b contains 4-8) while modeling how each strategy impacts the probability of above-market performance.
Let's cover both scenarios, starting with the one that aligns with your ranking goal.
Scenario 1: Rank Strategies by Impact on Performance (Hierarchical Logit)
We'll use the lme4 package to build a multilevel logit model, where we include random effects for strategy groups (a/b) to account for shared variation within each group, and fixed effects for individual strategies (to get their specific impact on performance).
Step 1: Prepare Your Data
First, make sure your variables are formatted correctly:
- Convert
DiffPerformanceHLto a binary numeric variable (1 for "H", 0 for "L") since logit models require numeric responses. - Ensure
StrategyandMainare coded as factors (to treat them as categorical variables).
# Load required packages library(lme4) library(dplyr) # Clean data your_data <- your_data %>% mutate( DiffPerformanceHL = ifelse(DiffPerformanceHL == "H", 1, 0), Strategy = as.factor(Strategy), Main = as.factor(Main) )
Step 2: Build the Hierarchical Logit Model
This model includes:
- Fixed effects for each strategy, control variables (
Control1,Control2) - Random intercept for strategy groups (
Main) to account for nested variation - Optional: Add a random intercept for
IDto control for firm-specific unobserved heterogeneity
# Hierarchical logit model with firm-level random effects hierarchical_logit <- glmer( DiffPerformanceHL ~ Strategy + Control1 + Control2 + (1 | ID) + (1 | Main), data = your_data, family = binomial(link = "logit") ) # View model summary summary(hierarchical_logit)
Step 3: Rank Strategies
To get the strategy ranking, we'll extract the fixed effect coefficients for each strategy, convert them to predicted probabilities (to make interpretation easier), then sort them:
# Extract strategy coefficients (excluding intercept and controls) strategy_coefs <- coef(summary(hierarchical_logit))[grepl("Strategy", rownames(coef(summary(hierarchical_logit)))), ] # Get model intercept intercept <- coef(summary(hierarchical_logit))["(Intercept)", "Estimate"] # Convert log-odds to predicted probability of above-market performance # (holding control variables at their mean values) strategy_probs <- data.frame( Strategy = gsub("Strategy", "", rownames(strategy_coefs)), LogOdds = strategy_coefs[, "Estimate"], Probability = exp(intercept + strategy_coefs[, "Estimate"]) / (1 + exp(intercept + strategy_coefs[, "Estimate"])) ) # Sort strategies by probability (highest to lowest performance likelihood) strategy_ranking <- strategy_probs %>% arrange(desc(Probability)) print(strategy_ranking)
Scenario 2: Nested Logit for Strategy Choice (If Your Goal Was Modeling Choice)
If you actually wanted to model why firms choose a specific strategy (with performance as a predictor), we'll use the mlogit package to build a nested logit model.
Step 1: Reshape Data for mlogit
mlogit requires data in "long" format where each row represents a possible strategy choice for a firm-month.
library(mlogit) # Reshape data to long format for discrete choice modeling long_data <- mlogit.data( your_data, choice = "Strategy", shape = "wide", alt.levels = as.character(1:8), chid.var = "ID" ) # Define the nesting structure (a group: 1-3, b group: 4-8) nesting <- list(a = c("1", "2", "3"), b = c("4", "5", "6", "7", "8"))
Step 2: Build the Nested Logit Model
nested_logit <- mlogit( Strategy ~ DiffPerformance + Control1 + Control2 | 0 | 0, data = long_data, nests = nesting, un.nest.el = FALSE ) summary(nested_logit)
Step 3: Interpret for Strategy Ranking
From this model, you can extract marginal effects of performance on choosing each strategy, then rank strategies by how sensitive their choice is to better performance (or directly compare predicted choice probabilities).
Handling Unbalanced Panel Data
Both models above naturally handle unbalanced panels:
glmeraccounts for repeated observations per firm via the(1 | ID)random effect, even if firms have differing numbers of months in the dataset.mlogitworks with unbalanced data as long as each firm has at least one observation of their chosen strategy.
内容的提问来源于stack exchange,提问作者Simon_W

