You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Statsmodels GLM泊松回归预测Quantity取值概率?

Understanding Poisson Regression Probability Predictions in Statsmodels

Great question—let’s unpack what’s happening here and clarify your confusion.

First, confirm your code logic is correct

Your approach to calculating the probability of Quantity taking specific integer values (1,2,3,4) is right on track. Here’s why:

  • When you fit a Poisson GLM with Statsmodels, the model outputs a predicted mean μ for each observation in your dataset (stored in series['poisson_predict']). This μ is determined by the unique combination of your independent variables (Month, cannibal_numbers, category_performance) for that row.
  • Using stats.poisson.pmf(counts, mu) computes the Poisson distribution’s probability mass function: for each observation’s μ, it calculates the probability that Quantity equals each value in counts.

Addressing your core question: "Does each demand value correspond to a unique probability?"

Short answer: No, not globally—but yes for each individual observation.

Let’s break this down:

  • The Poisson PMF depends entirely on the mean μ. Since your model generates a unique μ for every row (because each row has different values for your independent variables), the probability of Quantity=1 will vary from row to row. For example, a row with a higher cannibal_numbers might have a higher μ, so the probability of Quantity=3 would be different than a row with a lower μ.
  • For a single observation (fixed μ), each demand value (1,2,3,...n) does correspond to exactly one unique probability. That’s a defining property of the Poisson PMF—given a fixed mean, each integer k has a single, calculable probability.

Small optimization for your code

If you want probabilities for Quantity=1 through n (instead of starting at 0), adjust your counts definition. You can also name the result columns for clarity:

import numpy as np
import pandas as pd
import scipy.stats as stats
import statsmodels.api as sm
from statsmodels.formula.api import dmatrices

# Your existing model code (for context)
expr = """Quantity ~ Month + cannibal_numbers + category_performance"""
y, X = dmatrices(expr, series, return_type='dataframe')
poisson_fit = sm.GLM(y, X, family=sm.families.Poisson()).fit()
series['poisson_predict'] = poisson_fit.predict()

# Calculate probabilities for Quantity=1 to 4
n = 4
counts = np.arange(1, n+1)  # Starts at 1, ends at n
predict_prob = stats.poisson.pmf(counts, np.asarray(series['poisson_predict'])[:, None])
results = pd.DataFrame(predict_prob, columns=[f'P(Quantity={k})' for k in counts])

# Merge results back to original dataset
series = pd.concat([series, results], axis=1)

This gives you a dataframe where each row has columns like P(Quantity=1), P(Quantity=2), etc., showing the probability of that demand value for the specific combination of independent variables in that row.

内容的提问来源于stack exchange,提问作者Murcielago

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 07:59:36