You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

DataFrame数据提取与统计:运动员排名统计表格优化需求

问题与解决方案

原始数据与需求

已有如下DataFrame:

import pandas as pd

data = {'Eventname': ['100m','200m','Discus','100m','200m','Discus'],
        'Year': [2030,2030,2031,2030,2031,2032],
       'FirstPlace': ['John Smith', 'Shar jean', 'Abi whi', 'mik jon','joh doe', 'John Smith'],
        'SecPlace': ['joh doe', 'John Smith', 'Shar jean', 'Hen Hun','Tom Will', 'Gord Jay'],
        'thiPlace': ['mik jon', 'Lisa tru', 'John Smith', 'Bret Tun','Tim Smith', 'Jack Mann'] } 
df = pd.DataFrame(data)

需求:创建新DataFrame,第一列包含所有出现在FirstPlace、SecPlace、thiPlace中的不重复运动员姓名,后续列分别统计每个姓名在对应排名列中的出现次数。

原代码及问题

原编写代码:

NewArr=pd.DataFrame()
NewArr['first']=df['FirstPlace'].value_counts()
NewArr['second']=df['SecPlace'].value_counts()
NewArr['third']=df['thiPlace'].value_counts()

存在问题:

  • 仅显示FirstPlace列中的姓名,遗漏了仅出现在SecPlace或thiPlace中的运动员
  • 结果中存在NaN值,需要替换为0
  • 姓名列无标题,需添加“AthleteName”列名

修正后的代码

# 收集所有不重复的运动员姓名
all_athletes = pd.concat([df['FirstPlace'], df['SecPlace'], df['thiPlace']]).unique()

# 构建新DataFrame并处理所有需求
new_df = pd.DataFrame({
    'first': df['FirstPlace'].value_counts(),
    'second': df['SecPlace'].value_counts(),
    'third': df['thiPlace'].value_counts()
}).reindex(all_athletes, fill_value=0).reset_index().rename(columns={'index': 'AthleteName'})

代码说明

  1. 覆盖所有运动员:用pd.concat合并三个排名列,再通过unique()提取所有不重复姓名,确保没有遗漏仅出现在第二、第三名的运动员
  2. 替换NaN为0:使用reindex(all_athletes, fill_value=0),让每个排名的统计结果都对齐所有运动员,不存在的记录直接填充0
  3. 添加姓名列标题:通过reset_index()把原索引转为普通列,再用rename将列名改为“AthleteName”

内容的提问来源于stack exchange,提问作者almegdadi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 14:10:37