You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何识别MultiIndex层级0中对应Industry列不唯一的Ticker并处理重复计数

问题描述

我正在处理如下结构的面板数据(仅展示相关列):

Ticker Year Account_number Industry
AAA    2018  xxxx           Fossil
       2019  xxxx           Fossil
       2020  xxxx           Fossil
BBB    2018  yyyy           Materials
       2019  yyyy           Services
       2020  yyyy           Materials
CCC    2018  zzzz           Services
       2019  zzzz           Services
       2020  zzzz           Services

其中Ticker为MultiIndex的层级0,用于标识面板中的独立唯一单元,每个单元会被观测3年(MultiIndex的层级1)。执行groupby('Industry')操作时,部分Ticker(如'BBB')对应多个Industry,导致单元被重复计数。我需要识别并输出存在该问题的Ticker,以便手动为这些Ticker分配单一Industry。

解决方案

可以直接利用Pandas的分组和唯一值统计功能实现需求,代码如下:

# 假设你的DataFrame变量名为df
# 按MultiIndex的Ticker层级分组,统计每个Ticker对应的Industry唯一值数量
industry_unique_counts = df.groupby(level='Ticker')['Industry'].nunique()

# 筛选出Industry值不唯一的Ticker列表
problem_tickers = industry_unique_counts[industry_unique_counts > 1].index.tolist()

# 输出结果
print(problem_tickers)

代码说明

  1. groupby(level='Ticker'):直接基于MultiIndex的第0层(Ticker)进行分组,无需额外重置索引;
  2. nunique():统计每个分组内Industry列的唯一值数量;
  3. 筛选条件industry_unique_counts > 1:找出那些对应多个不同Industry的Ticker,最终转为列表输出。

运行上述代码后,会输出['BBB'],即所有存在Industry不一致问题的Ticker。

内容的提问来源于stack exchange,提问作者dm_

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 14:50:32