You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas多列值频率分析及高支持度问题筛选技术问询

问题描述

我有一份基于调查数据的Pandas DataFrame,包含24列(对应调查问题)与207行(对应受访者),列的取值(即答案)固定为:completely agree、rather agree、no idea、rather not agree、completely not agree。

我已经能用df['Q1'].value_counts()获取单个列(单个问题)的答案频率,但尝试多种方法获取多列的答案频率都没成功。

需求:

  1. 有没有办法在同一概览中查看多列的答案值频率?
  2. 找出支持度最高的3个问题(即“completely agree”与“rather agree”的频率之和),并按支持度从高到低排序。

示例数据集创建代码:

import pandas as pd
df = pd.read_csv('sampledata.csv')
df = pd.DataFrame({
'Q1':['Completely disagree','Completely disagree','Rather not agree','Rather agree'],
'Q2':['Completely disagree','Rather not agree','Rather agree','Rather agree'],
'Q3':['No idea / no opinion','Rather not agree','Rather agree','Rather agree'],
'Q4':['Completely disagree','Rather not agree','Completely disagree','Completely agree']
})
df.head()
解决方案

1. 多列答案频率概览

对DataFrame的每列单独执行value_counts,再用fillna(0)补全缺失的答案类型,转置后就能得到统一格式的多列频率概览:

# 计算每列答案频率,补全0值并转置以问题为行展示
freq_overview = df.apply(pd.Series.value_counts).fillna(0).astype(int).T
print(freq_overview)

对应测试数据的输出示例:

Completely disagree  Rather not agree  Rather agree  No idea / no opinion  Completely agree
Q1                     2                 1             1                     0                 0
Q2                     1                 1             2                     0                 0
Q3                     0                 1             2                     1                 0
Q4                     2                 1             0                     0                 1

2. 找出支持度最高的3个问题

先定义支持度为“completely agree”和“rather agree”的数量之和,对每列计算该值后排序取前3即可:

# 计算单列支持度的函数
def calculate_support(col):
    return col[(col == 'Completely agree') | (col == 'Rather agree')].count()

# 对所有列计算支持度,降序排序后取前3
support_ranking = df.apply(calculate_support).sort_values(ascending=False).head(3)
print(support_ranking)

对应测试数据的输出示例:

Q2    2
Q3    2
Q4    1
dtype: int64

如果需要查看支持度百分比(支持人数/总受访者),可修改函数为:

def calculate_support_percent(col):
    total = len(col)
    support_count = col[(col == 'Completely agree') | (col == 'Rather agree')].count()
    return round(support_count / total * 100, 2)

support_ranking_percent = df.apply(calculate_support_percent).sort_values(ascending=False).head(3)
print(support_ranking_percent)

对应测试数据的输出示例:

Q2    50.00
Q3    50.00
Q4    25.00
dtype: float64

内容的提问来源于stack exchange,提问作者Frank

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 18:45:50