使用Matplotlib构建调查响应报告:Likert分数高效统计及现有处理流程的优化咨询
使用Matplotlib构建调查响应报告:Likert分数高效统计及现有处理流程的优化咨询
我正在给一个现有应用开发功能,让用户上传Excel文件来可视化调查问题的响应数据。调查包含4个Likert量表问题和2个开放式评论问题,数据结构如下:
Q1 Q2 Q3 Q4 Q5 Q6 0 Occasionally Agree Agree comment comment Very Well 1 Frequently Agree Agree NaN NaN Moderately 2 Frequently Agree Agree NaN NaN Well 3 Frequently Agree Neither Agree nor Disagree comment comment Moderately
目前我通过以下代码来生成一个包含所有问题的列表,每个问题条目包含问题本身、响应选项、Likert量表类型,以及每个响应的出现次数:
def read_responses(xl): responses = [] action = False data = pd.read_excel(xl) for name, values in data.items(): if name.endswith("?"): action = True if action == True: data_set = {} set = values.value_counts(dropna=True).to_dict() labels, data = list(set.keys()), list(set.values()) data_set['question'] = name data_set['type'] = id_type(labels) data_set["responses"] = labels data_set["values"] = data responses.append(data_set) build_chart(responses)
这段代码会遍历DataFrame,跳过开头几列系统生成的无关信息,然后统计每列中各响应值的出现次数,整理成结构化数据后传给绘图函数用于后续可视化。
最终传递给绘图函数的数据集结构示例如下:
[{'question': 'Q1', 'type': 'likert', 'responses': ['Agree'], 'values': [4] }]
问题类型由以下函数判定:它会将响应选项的第一个值与预定义的Likert量表选项列表对比,如果该值不在任何一组Likert选项中,就判定为开放式评论问题:
def id_type(response): type = 'comment' for scale in scales: for term in scale: if term == response[0]: type = 'likert' return type
之后我会把这些数据传给一个用PdfPages和Matplotlib实现的绘图函数,生成PDF报告。我觉得当前的实现已经比较精简了,但还是想做个 sanity check——有没有更通用、更被行业广泛接受的调查响应数据处理方案?
备注:内容来源于stack exchange,提问作者Fennario
相关产品推荐
相关产品推荐

