You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

DataCamp诺贝尔项目报错:Series真值不明确,求解决方案

DataCamp诺贝尔项目报错分析与修复

问题描述

完成DataCamp诺贝尔项目时提交代码触发报错:

Your solution doesn't look quite right
The truth value of a Series is ambiguous. Use a.empty, a.bool(), a.item(), a.any() or a.all().

本地运行无明显问题,代码如下:

# Loading in required libraries
import pandas as pd
import seaborn as sns
import numpy as np

# Start coding here!

df = pd.read_csv('data/nobel.csv')

# most commonly awarded gender and birth country 
top_gender = df.sex.value_counts().index[0]
top_country = df.birth_country.value_counts().index[0]

# Calculate the proportion of USA born winners per decade
df['usa_born_winner'] = df['birth_country'] == 'United States of America'
df['decade'] = np.floor(df.year / 10) * 10
df['decade'] = df['decade'].astype(int)

# Identify the decade with the highest proportion of US-born winners
max_decade_usa = df.groupby('decade', as_index=False)['usa_born_winner'].mean()
max_decade_usa = max_decade_usa.loc[max_decade_usa['usa_born_winner'].idxmax(), 'decade']


# Calculating the proportion of female laureates per decade
max_female_dict = df.groupby(['decade', 'category'], as_index=False)['female_winner'].mean()

# Find the decade and category with the highest proportion of female laureates
max_female_dict = max_female_dict.loc[max_female_dict['female_winner'].idxmax(), ['decade', 'category']]
max_female_dict = {max_female_dict['decade']: max_female_dict['category']}


# Finding the first woman to win a Nobel Prize
woman = df[df['female_winner'] == True]
min_row = woman[woman['year'] == woman['year'].min()]
first_woman_name = min_row['full_name']
first_woman_category = min_row['category']

# Selecting the laureates that have received 2 or more prizes
repeats = df.full_name.value_counts()
repeats = repeats[repeats >= 2].index
repeat_list = list(repeats)

报错原因分析

报错核心是布尔Series的真值判断歧义,问题出在两处:

  1. 提取首位女性获奖者时,min_row = woman[woman['year'] == woman['year'].min()]返回DataFrame,后续first_woman_name和first_woman_category得到的是Series而非单个值。DataCamp判分系统尝试对这些Series进行布尔校验时,触发歧义报错。
  2. 构建max_female_dict时,max_female_dict.loc[..., ['decade', 'category']]返回的是Series,直接转为字典时,判分系统可能因类型解析异常触发错误。

本地运行无问题是因为Python允许直接输出Series,但DataCamp对变量类型有严格要求(需单个值而非Series)。

修复指导

针对上述问题,修改对应代码段:

1. 修复首位女性获奖者提取

将获取单个值的代码改为:

# Finding the first woman to win a Nobel Prize
woman = df[df['female_winner'] == True]
min_year = woman['year'].min()
min_row = woman[woman['year'] == min_year].iloc[0]  # 取第一行转为单个数据项
first_woman_name = min_row['full_name']  # 得到单个字符串
first_woman_category = min_row['category']  # 得到单个字符串

或更简洁的写法:

first_woman = df[df['female_winner'] == True].sort_values('year').iloc[0]
first_woman_name = first_woman['full_name']
first_woman_category = first_woman['category']

2. 修复最高女性占比的decade-category字典构建

确保从Series中提取单个值:

# Find the decade and category with the highest proportion of female laureates
max_female_row = max_female_dict.loc[max_female_dict['female_winner'].idxmax()]
max_female_dict = {int(max_female_row['decade']): max_female_row['category']}

用int()确保decade为整数类型,避免数据类型不一致触发判分错误。

3. 其他潜在优化(可选)

提取重复获奖者时,可改用更清晰的逻辑:

# Selecting the laureates that have received 2 or more prizes
repeat_list = df[df.duplicated('full_name', keep=False)]['full_name'].unique().tolist()

原代码无错误,此为可读性优化。

内容的提问来源于stack exchange,提问作者flydevilz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 05:52:47