如何计算缩放型调查数据的总计数百分比
CSV文件Answer字段占比统计实现方案
使用Python pandas库可快速生成你需要的两列统计结果,操作步骤如下:
- 安装依赖:如果没装pandas,先在终端执行
pip install pandas完成安装 - 读取本地CSV文件,提取
Answer字段做聚合统计 - 计算每个取值的出现次数占总样本量的比例,格式化为百分比
- 输出结果表,支持直接导出为新的CSV文件
核心实现代码:
import pandas as pd # 替换为你本地CSV文件的实际存储路径 df = pd.read_csv("你的调查数据文件路径.csv") # 统计每个Answer取值的出现次数 stat_df = df["Answer"].value_counts().sort_index().reset_index() stat_df.columns = ["Answer", "sample_count"] # 计算总计计数占比,默认保留2位小数,可按需调整round的参数 total_sample = len(df) stat_df["grand total count %"] = (stat_df["sample_count"] / total_sample * 100).round(2).astype(str) + "%" # 提取需要的两列作为最终结果 final_result = stat_df[["Answer", "grand total count %"]] # 可选:将结果保存为新CSV,utf-8-sig编码可避免Excel打开乱码 final_result.to_csv("answer占比统计结果.csv", index=False, encoding="utf-8-sig") # 控制台打印结果核对 print(final_result)
如果你的Answer字段是0-100区间的连续数值、需要按区间分箱统计(比如每10分为一个区间),可以将统计部分的代码替换为分箱逻辑:
# 按0-10、10-20...90-100的区间分箱,可修改bins参数调整间隔 df["answer_range"] = pd.cut(df["Answer"], bins=[0,10,20,30,40,50,60,70,80,90,100], include_lowest=True) stat_df = df["answer_range"].value_counts().sort_index().reset_index() stat_df.columns = ["Answer", "sample_count"]
内容的提问来源于stack exchange,提问作者Bhadriraju Nitin
相关产品推荐
相关产品推荐

