如何用Python Pandas编写循环追加DataFrame行的函数?
问题需求
遍历自定义的neighbourhood_group列表,将每个街区组中价格最低的房源数据追加到空DataFrame中。
示例数据集代码
import pandas as pd import numpy as np dict1 = {'id' : [2539,2595,3647,3831,12937,18198,258838,258876,267535,385824], 'name':['Clean & quiet apt home by the park','Skylit Midtown Castle','THE VILLAGE OF HARLEM....NEW YORK !','Cozy Entire Floor of Brownstone','1 Stop fr. Manhattan! Private Suite,Landmark Block','Little King of Queens','Oceanview,close to Manhattan','Affordable rooms,all transportation','Home Away From Home-Room in Bronx','New York City- Riverdale Modern two bedrooms unit'], 'price':[149,225,150,89,130,70,250,50,50,120], 'neighbourhood_group':['Brooklyn','Manhattan','Manhattan','Brooklyn','Queens','Queens','Staten Island','Staten Island','Bronx','Bronx']} df = pd.DataFrame(dict1)
原错误代码
nbd_grp = ['Bronx','Queens','Staten Islands','Brooklyn','Manhattan'] # 创建函数查找各街区组的最便宜房源 dfdf = pd.DataFrame(columns = ['id','name','price','neighbourhood_group']) def cheapest_place(neighbourhood_group): for elem in nbd_grp: data = df.loc[df['neighbourhood_group']==elem] cheapest = data.loc[data['price']==min(data['price'])] dfdf = cheapest.copy() cheapest_place(nbd_grp)
代码问题分析
- 拼写错误:
nbd_grp中的Staten Islands与数据集中的Staten Island不匹配,导致该街区组无法筛选出数据。 - 变量作用域问题:函数内部修改的
dfdf是局部变量,不会影响全局定义的空DataFrame,最终结果仍为空。 - 数据覆盖而非追加:每次循环用
cheapest.copy()直接覆盖dfdf,最后仅保留最后一个街区组的数据。 - 冗余逻辑:通过
min(data['price'])再筛选的方式可以简化,分组取最小值的方法更高效。
修正方案
方案一:修复原函数逻辑
nbd_grp = ['Bronx','Queens','Staten Island','Brooklyn','Manhattan'] # 创建空结果DataFrame dfdf = pd.DataFrame(columns = ['id','name','price','neighbourhood_group']) def cheapest_place(target_groups): global dfdf # 声明使用全局变量 for elem in target_groups: data = df.loc[df['neighbourhood_group'] == elem] if not data.empty: # 避免空数据引发的错误 # 筛选当前街区组价格最低的行 cheapest = data[data['price'] == data['price'].min()] # 追加到结果DataFrame,重置索引避免重复 dfdf = pd.concat([dfdf, cheapest], ignore_index=True) cheapest_place(nbd_grp) print(dfdf)
方案二:高效分组实现
如果无需严格匹配自定义列表顺序,可先分组取最小值再排序:
# 按街区组分组,取每组价格最低的行(若有多个最低价,取第一个出现的) result = df.loc[df.groupby('neighbourhood_group')['price'].idxmin()] # 按自定义列表重新排序 result = result.set_index('neighbourhood_group').reindex(nbd_grp).reset_index() print(result)
最终输出
| id | name | price | neighbourhood_group |
|---|---|---|---|
| 267535 | Home Away From Home-Room in Bronx | 50 | Bronx |
| 18198 | Little King of Queens | 70 | Queens |
| 258876 | Affordable rooms,all transportation | 50 | Staten Island |
| 3831 | Cozy Entire Floor of Brownstone | 89 | Brooklyn |
| 3647 | THE VILLAGE OF HARLEM....NEW YORK ! | 150 | Manhattan |
内容的提问来源于stack exchange,提问作者int_x
相关产品推荐
相关产品推荐

