在R的simstudy包中实现条件分类分布时遇报错求助
I’ve run into this exact issue with simstudy before! The error "Categorical distribution requires 2 or more probabilities" pops up because when defining a conditional categorical variable, every condition you specify needs to provide at least two probabilities (since categorical variables must have 2+ categories).
Let’s break down what went wrong and fix your code. First, here’s a minimal example that reproduces your error (matching your initial code structure):
library(simstudy) # Define non-random xNr def <- defData(varname = "xNr", dist = "nonrandom", formula = 7, id = "idnum") # This will throw the error: only one probability specified for the condition def <- defData(def, varname = "contributor", dist = "categorical", formula = "0.8 | xNr ==7") dat <- genData(100, def)
The problem here is that we’re only giving one probability (0.8) for the condition xNr ==7—but simstudy needs probabilities for all categories (at least two) to generate the categorical variable.
Fix 1: Use conditional probability vectors with ifelse
If you want contributor to have two categories (e.g., "Yes" and "No") with conditional probabilities based on xNr, you can pass a vector of probabilities in the formula using ifelse:
library(simstudy) def <- defData(varname = "xNr", dist = "nonrandom", formula =7, id = "idnum") # Define contributor with two categories, conditional on xNr def <- defData(def, varname = "contributor", dist = "categorical", formula = "c(0.7, 0.3)", # Probabilities for "Yes" and "No" when xNr=7 levels = c("Yes", "No")) # Generate data dat <- genData(100, def) # Check results table(dat$contributor)
If xNr had multiple values, you’d extend the ifelse to handle each case:
# Let's make xNr have two possible values first def <- defData(varname = "xNr", dist = "categorical", formula = "0.5;0.5", levels = c("7", "9")) # Conditional probabilities for each xNr value def <- defData(def, varname = "contributor", dist = "categorical", formula = "ifelse(xNr == '7', c(0.7, 0.3), c(0.2, 0.8))", levels = c("Yes", "No")) dat <- genData(100, def) table(dat$xNr, dat$contributor)
Fix 2: Use defDataCond for explicit conditional definitions
Another clean approach is to use defDataCond to define separate probability rules for each condition:
library(simstudy) def <- defData(varname = "xNr", dist = "nonrandom", formula =7, id = "idnum") # Define conditional rule for xNr=7 defCond <- defDataCond( def, varname = "contributor", dist = "categorical", formula = "0.7;0.3", # Probabilities for "Yes" and "No" levels = c("Yes", "No"), condition = "xNr ==7" ) dat <- genData(100, defCond)
This is especially useful if you have multiple complex conditions to handle—it keeps your code more readable.
The key takeaway: whenever you define a categorical variable (conditional or not), you must provide a set of probabilities for all categories (at least two) that sum to 1.
内容的提问来源于stack exchange,提问作者Monody

