如何通过循环批量创建基于指定行数据的Pandas新列
批量基于指定行数据创建新列的高效方法
假设我们有如下示例DataFrame:
import pandas as pd df = pd.DataFrame({ 'Country': ['China', 'Canada', 'Mexico', 'Ethiopia'], 'Population': [100, 15, 25, 30], 'Region': ['Asia', 'NAmerica', 'NAmerica', 'Africa'], 'HDI': ['High', 'V.High', 'Medium', 'Low'] })
需求是:基于埃塞俄比亚的Population、Region、HDI数据批量创建对应新列,让所有行的新列都填充埃塞俄比亚的对应值,替代手动逐个创建列的低效方式。
原方法的问题
之前手动逐个创建列的方式,在列数较多时非常繁琐且耗时:
df['Population_2'] = df['Population'][df['Country'] == "Ethiopia"] df['Region_2'] = df['Region'][df['Country'] == "Ethiopia"] df['Population_2'].fillna(method='ffill')
高效解决方案
方法1:循环批量创建列
先提取埃塞俄比亚的行数据,再遍历目标列批量生成新列:
# 获取埃塞俄比亚的行数据(转为Series) ethiopia_row = df[df['Country'] == 'Ethiopia'].iloc[0] # 定义需要复制的列名列表 target_cols = ['Population', 'Region', 'HDI'] # 循环创建新列 for col in target_cols: df[f'{col}_2'] = ethiopia_row[col]
方法2:向量式合并(无循环更高效)
直接提取埃塞俄比亚的行,复制为与原DataFrame行数一致的新表,重命名列后合并:
target_cols = ['Population', 'Region', 'HDI'] # 提取埃塞俄比亚的行,重命名列并重置索引 ethiopia_expanded = df[df['Country'] == 'Ethiopia'] \ .rename(columns={col: f'{col}_2' for col in target_cols}) \ .reset_index(drop=True) # 合并原表与扩展后的埃塞俄比亚数据 df = pd.concat([df, ethiopia_expanded], axis=1)
最终效果
执行后得到的DataFrame如下:
Country Population Region HDI Population_2 Region_2 HDI_2 0 China 100 Asia High 30 Africa Low 1 Canada 15 NAmerica V.High 30 Africa Low 2 Mexico 25 NAmerica Medium 30 Africa Low 3 Ethiopia 30 Africa Low 30 Africa Low
内容的提问来源于stack exchange,提问作者JodeCharger100
相关产品推荐
相关产品推荐

