如何编写函数实现Pandas Series按指定行数计算中位数?
按每n行计算中位数的函数实现
给定如下DataFrame:
import pandas as pd df = pd.DataFrame({ "value": [10,20,30,40,50,60,70,80,90,100] })
需要实现一个函数,接收pd.Series和参数n,计算每n行的中位数;当剩余行数不足n时,取剩余数据的中位数。
预期输出示例
- 当
n=2时,函数返回:
median 15 35 55 75 95
- 当
n=3时,函数返回:
median 20 50 80 100
参考函数(计算均值)
以下是一个实现按每n行计算均值的参考函数,我们可以在此基础上修改为计算中位数:
def find_mean(col, rows): """ col: pd.Series rows: 每一组的行数 """ if isinstance(col, pd.Series): col = col.to_numpy() mod = col.shape[0] % rows if mod != 0: exclude = col[-mod:] keep = col[: len(col) - mod] out = keep.reshape((int(keep.shape[0]/rows), int(rows))).mean(1) out = np.hstack((out, exclude.mean())) else: out = col.reshape((int(col.shape[0]/rows), int(rows))).mean(1) return out
修改后的中位数计算函数
只需将原函数中的mean替换为median即可,同时补充必要的导入:
import numpy as np import pandas as pd def find_median(col, rows): """ col: pd.Series - 待计算的序列 rows: int - 每组的行数 """ if isinstance(col, pd.Series): col = col.to_numpy() mod = col.shape[0] % rows if mod != 0: exclude = col[-mod:] keep = col[: len(col) - mod] # 对每组计算中位数 out = keep.reshape((int(keep.shape[0]/rows), rows)).median(axis=1) # 添加剩余部分的中位数 out = np.hstack((out, np.median(exclude))) else: out = col.reshape((int(col.shape[0]/rows), rows)).median(axis=1) # 转换为带索引的Series,匹配示例输出格式 return pd.Series(out, name='median')
验证示例
调用函数测试:
# 测试n=2 print(find_median(df['value'], 2)) # 输出: # 0 15 # 1 35 # 2 55 # 3 75 # 4 95 # Name: median, dtype: float64 # 测试n=3 print(find_median(df['value'], 3)) # 输出: # 0 20.0 # 1 50.0 # 2 80.0 # 3 100.0 # Name: median, dtype: float64
内容的提问来源于stack exchange,提问作者serdar_bay
相关产品推荐
相关产品推荐

