如何在Pandas中利用列值匹配对应列标题并生成新列
解决Pandas中通过year列匹配列标题生成新列的问题
需求说明
现有一个以itemName为索引的Pandas DataFrame,需根据year列的数值,匹配对应的年份列标题,生成新列year_name。示例数据及期望结果如下:
原始DataFrame
| itemName | 2020 | 2021 | 2022 | 2023 | 2024 | year |
|---|---|---|---|---|---|---|
| item1 | 5 | 20 | 10 | 10 | 50 | 3 |
| item2 | 10 | 10 | 50 | 20 | 40 | 2 |
| item3 | 12 | 35 | 73 | 10 | 54 | 4 |
期望结果
| itemName | 2020 | 2021 | 2022 | 2023 | 2024 | year | year_name |
|---|---|---|---|---|---|---|---|
| item1 | 5 | 20 | 10 | 10 | 50 | 3 | 2022 |
| item2 | 10 | 10 | 50 | 20 | 40 | 2 | 2021 |
| item3 | 12 | 35 | 73 | 10 | 54 | 4 | 2023 |
错误分析
你尝试的两段代码存在以下问题:
- 第一段代码报错
TypeError: list indices must be integers or slices, not Series:对单列使用apply时,lambda参数x是整个列的Series对象,而列表只能用整数/切片索引,无法直接用Series索引,因此报错。 - 第二段代码报错
IndexError: list index out of range:一是你依然按列处理apply,x.iloc[0]只会取第一行的year值,无法覆盖所有行;二是year列的数值是1-based索引(比如year=3对应第3个年份列),但列表是0-based索引,直接用原数值索引会超出列表范围。
修正方案
方法1:用apply逐行处理(修正lambda)
先提取年份列的名称列表,再对year列逐行应用lambda,将1-based的year值转换为0-based的列表索引:
# 提取年份列的名称列表(排除year列) col_names = [col for col in result_df.columns if col != 'year'] # 生成year_name列 result_df['year_name'] = result_df['year'].apply(lambda x: col_names[x - 1])
方法2:更简洁的map方式
如果year值和列表索引的对应关系固定(year值-1=列表索引),也可以用map实现:
col_names = [col for col in result_df.columns if col != 'year'] result_df['year_name'] = result_df['year'].map(lambda x: col_names[x-1])
执行上述代码后,即可得到你期望的DataFrame,year列的数值会正确匹配对应的年份列标题,生成year_name列。
内容的提问来源于stack exchange,提问作者DGMS89
相关产品推荐
相关产品推荐

