You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将DataFrame中每日为中心的5列垂直合并为新单列并计算90百分位

实现代码与思路说明

你原有代码的问题是通过行内apply把多列值拼接为字符串,属于水平聚合,要实现多列所有值的垂直堆叠,直接提取窗口列后拍平为一维数组即可,以下是可直接运行的实现:

方案1:先构造目标RollingPercentile再计算百分位

import pandas as pd
import numpy as np

# 先获取所有日期列名列表
date_cols = BaseResult.columns.tolist()
n_dates = len(date_cols)
half_window = 2 # 5天窗口中心前后各2天

rolling_data = {}
for target_idx in range(n_dates):
    # 生成当前目标列对应的窗口列索引,这里用取模实现边界循环(如1月1日的窗口会取上一年12月最后2天)
    window_cols_idx = [(target_idx - half_window + i) % n_dates for i in range(5)]
    # 提取5列所有值,拍平为一维数组(41行*5列=205个元素,你提到的200为笔误)
    window_vals = BaseResult.iloc[:, window_cols_idx].values.flatten()
    rolling_data[date_cols[target_idx]] = window_vals

# 得到你需要的205行×365列的RollingPercentile
RollingPercentile = pd.DataFrame(rolling_data)
# 计算每列90百分位,得到含365个值的数组
Percentile90_arr = RollingPercentile.quantile(0.9, axis=0).values

方案2:更省内存的直接计算方案

不需要构造中间的大体积RollingPercentile,遍历过程中直接计算百分位即可:

import numpy as np

date_cols = BaseResult.columns.tolist()
n_dates = len(date_cols)
half_window = 2
percentile90_list = []

for target_idx in range(n_dates):
    window_cols_idx = [(target_idx - half_window + i) % n_dates for i in range(5)]
    window_vals = BaseResult.iloc[:, window_cols_idx].values.flatten()
    percentile90_list.append(np.percentile(window_vals, 90))

# 直接得到365个值的结果数组
Percentile90_arr = np.array(percentile90_list)

边界处理说明

如果不需要循环取跨年日期,而是希望开头/结尾的窗口截断(如1月1日只取1月1日-1月3日的数据),把窗口索引生成代码替换为以下即可:

window_cols_idx = [max(0, min(target_idx - half_window + i, n_dates-1)) for i in range(5)]

内容的提问来源于stack exchange,提问作者Megan Martin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 05:06:05