You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为DataFrame添加continent列:列表元素重复7次后切换

问题:为DataFrame添加重复指定次数的洲名列

我需要给处理后的DataFrame添加名为continent的列,要求列表中的每个洲名重复7次(对应每个文件生成的DataFrame的行数)后,再切换到下一个洲名。

我尝试的代码

import numpy as np
frames = []
for file in files:
    df=wrangle(file)
    frames.append(df)
    continent = ["Central and South America", "Eurasia", "Africa", "Asia Pacific", "Europe", "Middle East", "North America"]
    arr = np.repeat(continent, len(df) // len(continent))
    #arr = np.concatenate([([x]) for x in continent], axis=0) 
    df['continent'] = pd.Series(arr, index=df.index[:len(arr)])
    
df = pd.concat(frames, ignore_index=True)
print(df.info())

得到的结果

Year    Coal    Oil Natural gas Other   MT CO2  continent
0   1990    58  422 104 NaN MT CO2  Central and South America
1   1995    62  501 125 NaN MT CO2  Eurasia
2   2000    79  577 171 NaN MT CO2  Africa
3   2005    80  614 218 NaN MT CO2  Asia Pacific
4   2010    99  723 270 NaN MT CO2  Europe
5   2015    132 777 305 NaN MT CO2  Middle East
6   2017    125 734 289 NaN MT CO2  North America
7   1990    899 777 1026    NaN MT CO2  Central and South America
8   1995    603 426 856 14.0    MT CO2  Eurasia

期望的结果

Year    Coal    Oil Natural gas Other   MT CO2  continent
0   1990    58  422 104 NaN MT CO2  Central and South America
1   1995    62  501 125 NaN MT CO2  Central and South America
2   2000    79  577 171 NaN MT CO2  Central and South America
3   2005    80  614 218 NaN MT CO2  Central and South America
4   2010    99  723 270 NaN MT CO2  Central and South America
5   2015    132 777 305 NaN MT CO2  Central and South America
6   2017    125 734 289 NaN MT CO2  Central and South America
7   1990    899 777 1026    NaN MT CO2  Eurasia
8   1995    603 426 856 14.0    MT CO2  Eurasia.......

解决方案

你的问题出在np.repeat的使用逻辑上:当前代码是让每个洲名重复len(df)//len(continent)次(这里刚好是1次),所以会生成一个每个洲名各出现一次的数组,导致列值轮流切换。

根据需求,每个文件生成的DataFrame对应一个洲名,且该洲名要填充整个DataFrame的continent列,直接对列赋值即可,Pandas会自动将单个值广播到所有行:

import pandas as pd
frames = []
continent_list = ["Central and South America", "Eurasia", "Africa", "Asia Pacific", "Europe", "Middle East", "North America"]

# 用enumerate获取当前文件的索引,对应洲名列表的位置
for idx, file in enumerate(files):
    df = wrangle(file)
    # 直接给列赋值对应的洲名,自动填充所有行
    df['continent'] = continent_list[idx]
    frames.append(df)

df = pd.concat(frames, ignore_index=True)
print(df.info())

如果你的files数量和continent_list长度不一致,可以用循环重复洲名列表,比如用itertools.cycle来循环取洲名:

from itertools import cycle
import pandas as pd

frames = []
continent_list = ["Central and South America", "Eurasia", "Africa", "Asia Pacific", "Europe", "Middle East", "North America"]
continent_cycle = cycle(continent_list)

for file in files:
    df = wrangle(file)
    df['continent'] = next(continent_cycle)
    frames.append(df)

df = pd.concat(frames, ignore_index=True)

这样就能实现每个洲名对应一个文件的DataFrame,所有行都填充该洲名的效果。


内容的提问来源于stack exchange,提问作者Stone Bee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 04:05:27