You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas中按百分比获取DataFrame的前若干行?

提取DataFrame前N%记录的实现方法

其实Pandas的head()方法本身并不支持直接传入百分比参数(也就是你想要的frac参数),不过我们可以用几种简单的方式实现你要的功能,甚至可以封装成类似head(frac=...)的写法。


基础实现:直接计算行数

最直接的思路就是先算出总行数的百分比,再把这个数值传给head():

import pandas as pd

# 假设你的DataFrame是df
df = pd.DataFrame({'col1': range(100)})

# 计算前5%的行数(用int取整,自动向下取整)
percent = 0.05
num_rows = int(len(df) * percent)

# 提取前N%的记录
top_percent = df.head(num_rows)

如果你的数据集很小,比如总行数是21,5%是1.05,int()会自动向下取整为1行。要是你想向上取整(哪怕只有0.1行也保留1条),可以用math.ceil():

import math
num_rows = math.ceil(len(df) * percent)

进阶封装:自定义类似head的方法

如果你想更接近df.head(frac=0.05)的写法,可以自己封装一个函数,或者给DataFrame添加自定义访问器:

方式1:简单函数封装

def head_frac(df, frac):
    num_rows = int(len(df) * frac)
    return df.head(num_rows)

# 使用时直接调用
top_5_percent = head_frac(df, 0.05)

方式2:自定义DataFrame访问器(更优雅)

这种方式可以让你像调用内置方法一样使用,避免函数名冲突:

import pandas as pd

@pd.api.extensions.register_dataframe_accessor("utils")
class DataFrameUtils:
    def __init__(self, df_obj):
        self._df = df_obj
    
    def head_frac(self, frac):
        num_rows = int(len(self._df) * frac)
        return self._df.head(num_rows)

# 使用示例
top_5_percent = df.utils.head_frac(0.05)

这样你就可以通过df.utils.head_frac(0.05)来提取前5%的记录,用法和内置方法非常相似。


内容的提问来源于stack exchange,提问作者Mohamed Thasin ah

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 11:00:32