如何用Python自动化实现方程组中的累积分布函数近似计算?
用Python自动化计算累积分布函数(CDF)近似值
我们需要根据观测数据,计算从t=0到数据最大值的所有整数t对应的CDF近似值,核心公式为:F(t) = (观测值≤t的数量) / 总观测数
实现方式
1. 纯Python实现(无依赖库)
适合小型数据集,无需额外安装库:
def compute_cdf(observations): total = len(observations) if not total: return {} max_val = max(observations) cdf_results = {} for t in range(0, max_val + 1): # 统计观测值≤t的数量 count = sum(1 for obs in observations if obs <= t) cdf_results[f"F({t})"] = count / total return cdf_results # 示例调用 list_goals = [1, 2, 2, 1, 2] cdf = compute_cdf(list_goals) # 输出结果 for key, value in cdf.items(): print(f"{key} = {value}")
运行后输出:
F(0) = 0.0 F(1) = 0.4 F(2) = 1.0
2. NumPy实现(高效处理大数据)
利用NumPy的向量化操作提升计算效率,适合大规模数据集:
import numpy as np def compute_cdf_numpy(observations): obs_array = np.array(observations) total = len(obs_array) if total == 0: return {} max_val = obs_array.max() t_values = np.arange(0, max_val + 1) # 批量计算每个t对应的观测值计数 counts = np.array([np.sum(obs_array <= t) for t in t_values]) cdf_results = {f"F({t})": count / total for t, count in zip(t_values, counts)} return cdf_results # 示例调用 list_goals = [1, 2, 2, 1, 2] cdf_np = compute_cdf_numpy(list_goals) for key, value in cdf_np.items(): print(f"{key} = {value}")
3. Pandas实现(便捷的数据分析流程)
适合需要整合到数据分析流水线的场景,利用Pandas的统计函数简化操作:
import pandas as pd def compute_cdf_pandas(observations): df = pd.DataFrame({"values": observations}) total = len(df) if total == 0: return {} max_val = df["values"].max() # 生成所有需要计算的t值 all_t = pd.Series(range(0, max_val + 1), name="t") # 统计每个值的出现次数,缺失值补0 count_series = df["values"].value_counts().reindex(all_t, fill_value=0) # 计算累积和并得到CDF值 cumsum = count_series.cumsum() cdf_results = {f"F({t})": cumsum[t] / total for t in all_t} return cdf_results # 示例调用 list_goals = [1, 2, 2, 1, 2] cdf_pd = compute_cdf_pandas(list_goals) for key, value in cdf_pd.items(): print(f"{key} = {value}")
内容的提问来源于stack exchange,提问作者Tappetinoorange
相关产品推荐
相关产品推荐

