如何基于Continent与pd.cut分箱生成带排序值计数的多级索引Series?
问题描述
现有如下数据集片段:
Country Continent % Renewable 0 China Asia (15.754, 29.228] 1 United States North America (2.213, 15.754] 2 Japan Asia (2.213, 15.754] 3 United Kingdom Europe (2.213, 15.754] 4 Russian Federation Europe (15.754, 29.228]
需要生成一个带多级索引的Series,按Continent分组,每个大洲下展示所有% Renewable分箱的国家数量(计数为0的分箱也要保留),示例输出如下:
Asia (2.213, 15.754] 3 (15.754, 29.228] 1 (29.228, 42.702] 2 (56.176, 69.65] 0 (42.702, 56.176] 0 Europe (2.213, 15.754] 2 (15.754, 29.228] 2 (29.228, 42.702] 0 (56.176, 69.65] 1 (42.702, 56.176] 0 >>....and so on
尝试代码:
groups = renew.groupby(['Continent', pd.cut(renew['% Renewable'], 5)])
报错信息:
TypeError: can only concatenate str (not "float") to str
需求总结:创建以Continent和% Renewable分箱为索引的Series,展示每个大洲各分箱的国家计数(含0值),且分箱需按数值顺序排列。
解决方案
错误原因
你的% Renewable列已经是分箱后的区间类型(或字符串格式的区间),而非原始数值列,再用pd.cut处理会触发类型错误——pd.cut要求输入数值,而你传入的是区间字符串/对象,导致字符串与浮点数拼接失败。
正确实现步骤
- 提取所有唯一分箱并按数值顺序排序,确保所有大洲使用同一套分箱标准
- 用交叉表统计大洲与分箱的组合计数,自动填充缺失分箱的0值
- 转换为多级索引的Series
代码实现
import pandas as pd # 假设数据框名为renew # 1. 提取所有分箱并按区间左边界排序,保证分箱顺序正确 sorted_bins = sorted(renew['% Renewable'].unique(), key=lambda x: x.left) # 2. 生成交叉表,自动填充所有分箱的计数(无数据则为0) cross_table = pd.crosstab(renew['Continent'], renew['% Renewable'], dropna=False) # 3. 按排序后的分箱重新排列列,确保输出顺序符合要求 cross_table = cross_table[sorted_bins] # 4. 转换为多级索引的Series result = cross_table.stack()
补充:若% Renewable是字符串格式的区间
如果你的分箱是以字符串形式存储(而非pd.Interval对象),先把它转回区间类型:
# 将字符串区间转为pd.Interval对象 renew['% Renewable'] = renew['% Renewable'].apply(pd.Interval) # 再执行上述统计步骤
内容的提问来源于stack exchange,提问作者Brian Cox
相关产品推荐
相关产品推荐

