DataCamp诺贝尔项目报错:Series真值不明确,求解决方案
DataCamp诺贝尔项目报错分析与修复
问题描述
完成DataCamp诺贝尔项目时提交代码触发报错:
Your solution doesn't look quite right The truth value of a Series is ambiguous. Use a.empty, a.bool(), a.item(), a.any() or a.all().
本地运行无明显问题,代码如下:
# Loading in required libraries import pandas as pd import seaborn as sns import numpy as np # Start coding here! df = pd.read_csv('data/nobel.csv') # most commonly awarded gender and birth country top_gender = df.sex.value_counts().index[0] top_country = df.birth_country.value_counts().index[0] # Calculate the proportion of USA born winners per decade df['usa_born_winner'] = df['birth_country'] == 'United States of America' df['decade'] = np.floor(df.year / 10) * 10 df['decade'] = df['decade'].astype(int) # Identify the decade with the highest proportion of US-born winners max_decade_usa = df.groupby('decade', as_index=False)['usa_born_winner'].mean() max_decade_usa = max_decade_usa.loc[max_decade_usa['usa_born_winner'].idxmax(), 'decade'] # Calculating the proportion of female laureates per decade max_female_dict = df.groupby(['decade', 'category'], as_index=False)['female_winner'].mean() # Find the decade and category with the highest proportion of female laureates max_female_dict = max_female_dict.loc[max_female_dict['female_winner'].idxmax(), ['decade', 'category']] max_female_dict = {max_female_dict['decade']: max_female_dict['category']} # Finding the first woman to win a Nobel Prize woman = df[df['female_winner'] == True] min_row = woman[woman['year'] == woman['year'].min()] first_woman_name = min_row['full_name'] first_woman_category = min_row['category'] # Selecting the laureates that have received 2 or more prizes repeats = df.full_name.value_counts() repeats = repeats[repeats >= 2].index repeat_list = list(repeats)
报错原因分析
报错核心是布尔Series的真值判断歧义,问题出在两处:
- 提取首位女性获奖者时,
min_row = woman[woman['year'] == woman['year'].min()]返回DataFrame,后续first_woman_name和first_woman_category得到的是Series而非单个值。DataCamp判分系统尝试对这些Series进行布尔校验时,触发歧义报错。 - 构建
max_female_dict时,max_female_dict.loc[..., ['decade', 'category']]返回的是Series,直接转为字典时,判分系统可能因类型解析异常触发错误。
本地运行无问题是因为Python允许直接输出Series,但DataCamp对变量类型有严格要求(需单个值而非Series)。
修复指导
针对上述问题,修改对应代码段:
1. 修复首位女性获奖者提取
将获取单个值的代码改为:
# Finding the first woman to win a Nobel Prize woman = df[df['female_winner'] == True] min_year = woman['year'].min() min_row = woman[woman['year'] == min_year].iloc[0] # 取第一行转为单个数据项 first_woman_name = min_row['full_name'] # 得到单个字符串 first_woman_category = min_row['category'] # 得到单个字符串
或更简洁的写法:
first_woman = df[df['female_winner'] == True].sort_values('year').iloc[0] first_woman_name = first_woman['full_name'] first_woman_category = first_woman['category']
2. 修复最高女性占比的decade-category字典构建
确保从Series中提取单个值:
# Find the decade and category with the highest proportion of female laureates max_female_row = max_female_dict.loc[max_female_dict['female_winner'].idxmax()] max_female_dict = {int(max_female_row['decade']): max_female_row['category']}
用int()确保decade为整数类型,避免数据类型不一致触发判分错误。
3. 其他潜在优化(可选)
提取重复获奖者时,可改用更清晰的逻辑:
# Selecting the laureates that have received 2 or more prizes repeat_list = df[df.duplicated('full_name', keep=False)]['full_name'].unique().tolist()
原代码无错误,此为可读性优化。
内容的提问来源于stack exchange,提问作者flydevilz
相关产品推荐
相关产品推荐

