You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于DataFrame列索引值,从两列表生成计数新DataFrame

问题描述

我有两个列表:

a = [12, 12, 12, 3, 4, 5]
b = [1, 2, 4, 5, 6, 12, 4, 7, 9, 2, 3, 5, 6]

还有一个DataFrame df,它的列位置索引对应上述列表中的元素,比如df的列信息如下:

df.columns
Index(['lep', 'eta', 'phi', 'missing energy magn', 'missing energy phi', 'jet'])

(实际列数比示例更多)

我需要创建一个新的DataFrame,包含以下列:
索引值, 列名, 在a中的出现次数, 在b中的出现次数

举个例子:如果索引值12在列表a中出现3次、在列表b中出现2次,对应df中的列名为'foo',那么新DataFrame里的对应行就是:12, foo, 3, 2。

我知道可以用循环统计列表元素的计数,但不知道怎么把计数和对应的列索引关联起来。期望的输出示例如下:

new_df.head()
   index                name  count_a  count_b
0      0                 lep        0        0
1      1                 eta        0        1
2      2                 phi        0        2
3      3  missing energy magn        1        1
4      4  missing energy phi        1        1
解决方案

可以用collections.Counter快速统计列表元素的出现次数,再结合Pandas的映射功能关联到df的列信息,无需循环即可实现:

  1. 导入所需库
import pandas as pd
from collections import Counter
  1. 统计列表a和b中各索引值的出现次数
count_a = Counter(a)
count_b = Counter(b)
  1. 基于df的列信息创建基础DataFrame
# 生成列的位置索引和对应列名的数据集
base_data = {
    'index': range(len(df.columns)),
    'name': df.columns.tolist()
}
new_df = pd.DataFrame(base_data)
  1. 映射计数到新DataFrame中
# 用map匹配索引值的计数,不存在的索引值填充0并转为整数类型
new_df['count_a'] = new_df['index'].map(count_a).fillna(0).astype(int)
new_df['count_b'] = new_df['index'].map(count_b).fillna(0).astype(int)

执行完以上代码后,new_df就是符合需求的结果。


内容的提问来源于stack exchange,提问作者PolarVortex8

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 14:01:39