You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于参数n筛选DataFrame并保留同成本行的技术实现问题

问题描述

我有一个包含各类产品及其对应成本的DataFrame,现有一个通过参数n筛选DataFrame行的函数。例如当n=6时,会返回6行数据,但这样会遗漏第7行(该行可能与第6行成本相同)。我希望在基于n筛选的同时,返回所有与筛选结果中最后一行成本相同的额外行。目前我写出了如下代码,当n=4时它能返回5行(n=4加上1行成本相同的行),但无法返回第6行(若其成本也相同)。我是新手,希望能得到技术帮助!

def most_expensive_items(n): 
    subsetted_cost=COST_DF.sort_values(by=["cost"], ascending=False)[:n]
    test1_df=COST_DF.sort_values(by=["cost"], ascending=False)[:n+1]
    test2_df=COST_DF.sort_values(by=["cost"], ascending=False)[:n+2]
    test3_df=COST_DF.sort_values(by=["cost"], ascending=False)[:n+3]
    test4_df=COST_DF.sort_values(by=["cost"], ascending=False)[:n+4]
    test5_df=COST_DF.sort_values(by=["cost"], ascending=False)[:n+5]

    for r in COST_DF:
        for _, row_q in test1_df.iterrows():
            cost1= (test1_df.iloc[-1])
            for _, row_s in test2_df.iterrows():
                cost2= (test2_df.iloc[-1])
                for _, row_p in test3_df.iterrows():
                    cost3= (test3_df.iloc[-1])
                    for _, row_t in subsetted_cost.iterrows():
                        cost_last_row= (subsetted_cost.iloc[-1])
                    
        if cost1.cost == cost_last_row.cost:
            return subsetted_cost.append(cost1)
        if cost2.cost == cost_last_row.cost:
            return subsetted_cost.append(cost1, cost2)
        if cost3.cost == cost_last_row.cost:
            return subsetted_cost.append(cost1, cost2, cost3)
            
most_expensive_items(4)
解决方案

你的代码问题在于用固定的n+1到n+5范围判断,没法覆盖所有成本相同的行,而且嵌套循环冗余且效率极低。正确的思路是通过阈值筛选来实现需求,步骤如下:

  1. 先将整个DataFrame按成本降序排序
  2. 提取前n行中的最低成本作为阈值
  3. 筛选所有成本大于等于该阈值的行

优化后的代码:

def most_expensive_items(n):
    # 按成本降序排序,重置索引避免原索引混乱
    sorted_df = COST_DF.sort_values(by="cost", ascending=False).reset_index(drop=True)
    # 获取前n行的最低成本(索引从0开始,第n行对应索引n-1)
    threshold_cost = sorted_df.iloc[n-1]["cost"]
    # 筛选所有成本不低于阈值的行
    result = sorted_df[sorted_df["cost"] >= threshold_cost]
    return result

代码说明

  • sort_values完成降序排序,reset_index(drop=True)重置索引,避免原索引干扰后续筛选
  • 用iloc[n-1]精准定位前n行的最后一行,取其成本作为筛选阈值
  • 布尔索引sorted_df["cost"] >= threshold_cost会自动匹配所有成本符合要求的行,不管有多少行成本相同,都能全部返回

比如当n=4时,如果第4、5、6行成本一致,这个函数会直接返回前6行,完全覆盖所有符合条件的数据。

内容的提问来源于stack exchange,提问作者user20473020

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 11:01:59