如何基于Counter的keys和values创建Pandas DataFrame
解决从单列数据生成含唯一变量与计数的DataFrame问题
问题描述
想要从一列数据生成一个DataFrame,其中一列包含每个唯一变量,另一列为对应计数。尝试两种方法均未得到预期结果:
第一种尝试代码
def site_occurence(countycolumn): print('Site Occurrence') countylist = [] for county in countycolumn: countylist.append(county) a = Counter(countylist).keys() b = Counter(countylist).values() site_occurence = dict(zip(a, b)) df = pd.DataFrame(site_occurence) return df
终端输出:
Site Occurrence site occurrence 1685
第二种尝试代码
def site_occurence(countycolumn): print('Site Occurrence') countylist = [] for county in countycolumn: countylist.append(county) a = Counter(countylist).keys() b = Counter(countylist).values() df = pd.DataFrame((a, b), columns = ['site', 'occurrence']) return df
终端输出:
Site Occurrence 0 1 1685
预期输出格式:
Angus 11 Angus / East Perthshire 1 Argyll 176 Ayrshire 62 Banffshire 7 Banffshire / Moray ...........
遇到的问题:返回dict_keys和dict_values对象后,无法正确转换为符合预期的DataFrame。
错误原因分析
- 第一种方法中,直接用字典创建DataFrame时,pandas会将字典的键作为列名,值作为列数据,导致输出格式与预期完全相反。
- 第二种方法中,传入
(a, b)作为数据源时,pandas会将这两个对象视为两行数据,而指定的列数与每行元素数量不匹配,最终输出混乱。
正确实现方法
方法1:修正Counter转换逻辑
将dict_keys和dict_values转换为列表,构造(地区, 计数)的二维结构后传入DataFrame:
from collections import Counter import pandas as pd def site_occurence(countycolumn): print('Site Occurrence') countylist = list(countycolumn) # 无需手动循环追加,直接转列表 counter = Counter(countylist) # 构造符合DataFrame要求的二维数据结构 data = list(counter.items()) df = pd.DataFrame(data, columns=['site', 'occurrence']) return df
方法2:使用pandas内置高效方法
pandas自带的value_counts()方法可直接对列做计数,转换为DataFrame并重命名列即可,无需手动调用Counter:
import pandas as pd def site_occurence(countycolumn): print('Site Occurrence') # value_counts()返回计数结果的Series,转成DataFrame后重命名列 df = countycolumn.value_counts().reset_index() df.columns = ['site', 'occurrence'] return df
该方法更简洁,且处理大数据时效率更高,因为pandas内置方法经过性能优化。
验证结果
调用上述任一方法后,输出的DataFrame将符合预期格式:
| site | occurrence |
|---|---|
| Angus | 11 |
| Angus / East Perthshire | 1 |
| Argyll | 176 |
| Ayrshire | 62 |
| Banffshire | 7 |
内容的提问来源于stack exchange,提问作者Daisy
相关产品推荐
相关产品推荐

