如何用Statsmodels GLM泊松回归预测Quantity取值概率?
Great question—let’s unpack what’s happening here and clarify your confusion.
First, confirm your code logic is correct
Your approach to calculating the probability of Quantity taking specific integer values (1,2,3,4) is right on track. Here’s why:
- When you fit a Poisson GLM with Statsmodels, the model outputs a predicted mean μ for each observation in your dataset (stored in
series['poisson_predict']). This μ is determined by the unique combination of your independent variables (Month,cannibal_numbers,category_performance) for that row. - Using
stats.poisson.pmf(counts, mu)computes the Poisson distribution’s probability mass function: for each observation’s μ, it calculates the probability thatQuantityequals each value incounts.
Addressing your core question: "Does each demand value correspond to a unique probability?"
Short answer: No, not globally—but yes for each individual observation.
Let’s break this down:
- The Poisson PMF depends entirely on the mean μ. Since your model generates a unique μ for every row (because each row has different values for your independent variables), the probability of
Quantity=1will vary from row to row. For example, a row with a highercannibal_numbersmight have a higher μ, so the probability ofQuantity=3would be different than a row with a lower μ. - For a single observation (fixed μ), each demand value (1,2,3,...n) does correspond to exactly one unique probability. That’s a defining property of the Poisson PMF—given a fixed mean, each integer k has a single, calculable probability.
Small optimization for your code
If you want probabilities for Quantity=1 through n (instead of starting at 0), adjust your counts definition. You can also name the result columns for clarity:
import numpy as np import pandas as pd import scipy.stats as stats import statsmodels.api as sm from statsmodels.formula.api import dmatrices # Your existing model code (for context) expr = """Quantity ~ Month + cannibal_numbers + category_performance""" y, X = dmatrices(expr, series, return_type='dataframe') poisson_fit = sm.GLM(y, X, family=sm.families.Poisson()).fit() series['poisson_predict'] = poisson_fit.predict() # Calculate probabilities for Quantity=1 to 4 n = 4 counts = np.arange(1, n+1) # Starts at 1, ends at n predict_prob = stats.poisson.pmf(counts, np.asarray(series['poisson_predict'])[:, None]) results = pd.DataFrame(predict_prob, columns=[f'P(Quantity={k})' for k in counts]) # Merge results back to original dataset series = pd.concat([series, results], axis=1)
This gives you a dataframe where each row has columns like P(Quantity=1), P(Quantity=2), etc., showing the probability of that demand value for the specific combination of independent variables in that row.
内容的提问来源于stack exchange,提问作者Murcielago

