如何在Python Pandas中拆分列并按地区归类输出国家?
Pandas列拆分与国家地区归类问题
问题描述
我是Pandas新手,需要将数据列中的特定部分拆分,把国家按亚洲、欧洲、其他地区归类。尝试了代码但未成功,相关信息如下:
初始代码
import pandas as pd df = pd.read_csv('countries.csv') countrieslist = { 'Asia': list(df.columns.values), 'Europe': list(df.columns.values), 'Others': list(df.columns.values) } print(f"Countries in Asia - {countrieslist['Asia']}") print(f"Countries in Europe - {countrieslist['Europe']}") print(f"Countries in Others - {countrieslist['Others']}")
当前错误输出
Countries in Asia - [' ', ' Brunei Darussalam ', ' Indonesia ', ' Malaysia ', ' Philippines ', ' Thailand ', ' Viet Nam ', ' Myanmar ', ' Japan ', ' Hong Kong ', ' China ', ' Taiwan ', ' Korea, Republic Of ', ' India ', ' Pakistan ', ' Sri Lanka ', ' Saudi Arabia ', ' Kuwait ', ' UAE ', ' United Kingdom ', ' Germany ', ' France ', ' Italy ', ' Netherlands ', ' Greece ', ' Belgium & Luxembourg ', ' Switzerland ', ' Austria ', ' Scandinavia ', ' CIS & Eastern Europe ', ' USA ', ' Canada ', ' Australia ', ' New Zealand ', ' Africa ']
期望输出
Countries in Asia - ' Brunei Darussalam ', ' Indonesia ', ' Malaysia ', ' Philippines ', ' Thailand ', ' Viet Nam ', ' Myanmar ', ' Japan ', ' Hong Kong ', ' China ', ' Taiwan ', ' Korea, Republic Of ', ' India ', ' Pakistan ', ' Sri Lanka ', ' Saudi Arabia ', ' Kuwait ', ' UAE ' Countries in Europe - ' United Kingdom ', ' Germany ', ' France ', ' Italy ', ' Netherlands ', ' Greece ', ' Belgium & Luxembourg ', ' Switzerland ', ' Austria ', ' Scandinavia ', ' CIS & Eastern Europe ' Countries in Others – ' USA ', ' Canada ', ' Australia ', ' New Zealand ', ' Africa '
补充信息
print(df.columns)输出的列名列表包含一个空字符串列,后续依次是亚洲国家、欧洲国家、其他地区的名称(与错误输出中的列名顺序一致)。
解决方案
方法1:按列名位置切片(适合列顺序固定的情况)
先清洗掉空列,再按已知的列顺序拆分归类:
import pandas as pd df = pd.read_csv('countries.csv') # 清洗列名:去掉空的列(全是空格的列) cleaned_cols = [col for col in df.columns if col.strip()] # 按位置拆分对应地区(根据你的列顺序) countries_list = { 'Asia': cleaned_cols[:18], # 前18个是亚洲国家 'Europe': cleaned_cols[18:29], # 接下来11个是欧洲国家 'Others': cleaned_cols[29:] # 剩余5个是其他地区 } # 按期望格式打印 print(f"Countries in Asia - {', '.join(countries_list['Asia'])}") print(f"Countries in Europe - {', '.join(countries_list['Europe'])}") print(f"Countries in Others – {', '.join(countries_list['Others'])}")
方法2:按国家集合匹配(鲁棒性更强,不受列顺序影响)
手动定义各地区的国家/地区集合,通过匹配来归类:
import pandas as pd df = pd.read_csv('countries.csv') # 定义各地区的国家/地区集合(注意去掉前后空格) asia_set = { 'Brunei Darussalam', 'Indonesia', 'Malaysia', 'Philippines', 'Thailand', 'Viet Nam', 'Myanmar', 'Japan', 'Hong Kong', 'China', 'Taiwan', 'Korea, Republic Of', 'India', 'Pakistan', 'Sri Lanka', 'Saudi Arabia', 'Kuwait', 'UAE' } europe_set = { 'United Kingdom', 'Germany', 'France', 'Italy', 'Netherlands', 'Greece', 'Belgium & Luxembourg', 'Switzerland', 'Austria', 'Scandinavia', 'CIS & Eastern Europe' } others_set = { 'USA', 'Canada', 'Australia', 'New Zealand', 'Africa' } # 清洗列名:去掉空列,并保留带空格的原始格式 cleaned_cols = [col for col in df.columns if col.strip()] # 匹配归类 countries_list = { 'Asia': [col for col in cleaned_cols if col.strip() in asia_set], 'Europe': [col for col in cleaned_cols if col.strip() in europe_set], 'Others': [col for col in cleaned_cols if col.strip() in others_set] } # 按期望格式打印 print(f"Countries in Asia - {', '.join(countries_list['Asia'])}") print(f"Countries in Europe - {', '.join(countries_list['Europe'])}") print(f"Countries in Others – {', '.join(countries_list['Others'])}")
说明
- 两种方法都先过滤了空列(即那个全是空格的列),避免干扰归类。
- 方法1适合列顺序固定的场景,代码更简洁;方法2不受列顺序影响,即使数据列顺序调整也能正确归类。
内容的提问来源于stack exchange,提问作者sfrben99
相关产品推荐
相关产品推荐

