如何为DataFrame添加continent列:列表元素重复7次后切换
问题:为DataFrame添加重复指定次数的洲名列
我需要给处理后的DataFrame添加名为continent的列,要求列表中的每个洲名重复7次(对应每个文件生成的DataFrame的行数)后,再切换到下一个洲名。
我尝试的代码
import numpy as np frames = [] for file in files: df=wrangle(file) frames.append(df) continent = ["Central and South America", "Eurasia", "Africa", "Asia Pacific", "Europe", "Middle East", "North America"] arr = np.repeat(continent, len(df) // len(continent)) #arr = np.concatenate([([x]) for x in continent], axis=0) df['continent'] = pd.Series(arr, index=df.index[:len(arr)]) df = pd.concat(frames, ignore_index=True) print(df.info())
得到的结果
Year Coal Oil Natural gas Other MT CO2 continent 0 1990 58 422 104 NaN MT CO2 Central and South America 1 1995 62 501 125 NaN MT CO2 Eurasia 2 2000 79 577 171 NaN MT CO2 Africa 3 2005 80 614 218 NaN MT CO2 Asia Pacific 4 2010 99 723 270 NaN MT CO2 Europe 5 2015 132 777 305 NaN MT CO2 Middle East 6 2017 125 734 289 NaN MT CO2 North America 7 1990 899 777 1026 NaN MT CO2 Central and South America 8 1995 603 426 856 14.0 MT CO2 Eurasia
期望的结果
Year Coal Oil Natural gas Other MT CO2 continent 0 1990 58 422 104 NaN MT CO2 Central and South America 1 1995 62 501 125 NaN MT CO2 Central and South America 2 2000 79 577 171 NaN MT CO2 Central and South America 3 2005 80 614 218 NaN MT CO2 Central and South America 4 2010 99 723 270 NaN MT CO2 Central and South America 5 2015 132 777 305 NaN MT CO2 Central and South America 6 2017 125 734 289 NaN MT CO2 Central and South America 7 1990 899 777 1026 NaN MT CO2 Eurasia 8 1995 603 426 856 14.0 MT CO2 Eurasia.......
解决方案
你的问题出在np.repeat的使用逻辑上:当前代码是让每个洲名重复len(df)//len(continent)次(这里刚好是1次),所以会生成一个每个洲名各出现一次的数组,导致列值轮流切换。
根据需求,每个文件生成的DataFrame对应一个洲名,且该洲名要填充整个DataFrame的continent列,直接对列赋值即可,Pandas会自动将单个值广播到所有行:
import pandas as pd frames = [] continent_list = ["Central and South America", "Eurasia", "Africa", "Asia Pacific", "Europe", "Middle East", "North America"] # 用enumerate获取当前文件的索引,对应洲名列表的位置 for idx, file in enumerate(files): df = wrangle(file) # 直接给列赋值对应的洲名,自动填充所有行 df['continent'] = continent_list[idx] frames.append(df) df = pd.concat(frames, ignore_index=True) print(df.info())
如果你的files数量和continent_list长度不一致,可以用循环重复洲名列表,比如用itertools.cycle来循环取洲名:
from itertools import cycle import pandas as pd frames = [] continent_list = ["Central and South America", "Eurasia", "Africa", "Asia Pacific", "Europe", "Middle East", "North America"] continent_cycle = cycle(continent_list) for file in files: df = wrangle(file) df['continent'] = next(continent_cycle) frames.append(df) df = pd.concat(frames, ignore_index=True)
这样就能实现每个洲名对应一个文件的DataFrame,所有行都填充该洲名的效果。
内容的提问来源于stack exchange,提问作者Stone Bee
相关产品推荐
相关产品推荐

