如何在Pandas DataFrame中使用apply方法实现组织分类需求
解决方案
你遇到的问题主要是两个点:一是apply默认按列处理,没指定axis=1;二是错误用row.index获取组织名称,应该用row.name。下面给你几种可行的方案:
方案1:修正你的apply函数
def classify_org(row): if row['NumberofSpaceMission'] < 10: return 'Other organisation' return row.name # 必须指定axis=1,让apply逐行处理 new_df['important_organisations'] = new_df.apply(classify_org, axis=1)
方案2:用numpy.where(推荐,更高效)
对于DataFrame的条件赋值,用numpy.where比逐行apply快很多,尤其是数据量大的时候:
import numpy as np new_df['important_organisations'] = np.where( new_df['NumberofSpaceMission'] < 10, 'Other organisation', new_df.index )
方案3:用loc索引赋值
先把新列全部设为组织名称(即索引),再把符合条件的行替换成指定字符串:
new_df['important_organisations'] = new_df.index new_df.loc[new_df['NumberofSpaceMission'] < 10, 'important_organisations'] = 'Other organisation'
原代码失败的原因
- 缺少
axis=1参数:apply默认对每一列操作,函数接收的是列的Series而非行,所以row.index返回的是所有列名,不是你要的组织名称。 - 错误使用
row.index:即使按行处理,row是代表一行的Series,row.index是这行的列名列表,要获取当前行的索引(组织名)得用row.name。
内容的提问来源于stack exchange,提问作者Jossoft
相关产品推荐
相关产品推荐

