三款同类软件对比测试:用户样本与Likert量表调研咨询
Hey there! Let's work through your comparative software testing challenge step by step. With your small, partially overlapping user sample and Likert-scale questions, we can still run a rigorous, actionable analysis—here's how to approach it:
Your 30 users fall into three distinct buckets that you need to separate for analysis to avoid skewed results:
- Group A (15 users): Only used Software 1
- Group B (5 users): Only used Software 2
- Group C (5 users): Used both Software 1 and Software 2
Group C is especially valuable because they can directly compare the two existing tools—their paired feedback will be more reliable than cross-group comparisons between A and B.
Likert-scale scores are ordinal data (not interval/ratio), so skip parametric tests like t-tests. Instead, use these non-parametric methods tailored to small samples:
- Cross-group comparisons (e.g., Group A's rating of Software 1 vs Group B's rating of Software 2): Use the Mann-Whitney U Test—it's designed for independent groups with small, non-normal datasets.
- Paired comparisons (e.g., Group C's rating of Software 1 vs their rating of Software 2): Use the Wilcoxon Signed-Rank Test—this accounts for the same user providing both sets of scores, reducing variability.
- If you're including your own software in the ratings: For Group A, compare their scores of your software vs Software 1 (Wilcoxon, since it's paired for their experience); same for Group B with your software vs Software 2. Group C can compare all three, but you'll need to use a Friedman Test for paired ordinal data across three groups.
Don't aggregate scores across your 4 dimensions—each one (e.g., usability, feature completeness, performance) tells a unique story. For each dimension:
- Report median and interquartile range (IQR) instead of mean (median is more representative for ordinal data). You can include the mean as supplementary info, but prioritize median/IQR.
- Visualize results with boxplots—they'll make it easy to spot differences in score distributions between software tools and user groups.
Your sample is small (especially Group B with only 5 users), so be transparent about this in your findings:
- Frame your results as exploratory rather than definitive—they can't be generalized to all users of these tools, but they'll give you directional insights.
- Calculate effect sizes to quantify how large the differences are (e.g., for Mann-Whitney, use the r-value:
r = U / (n1*n2)). A large effect size even with a small sample means the difference is meaningful, not just random.
- Data organization: Use a spreadsheet (Excel/Google Sheets) or statistical tool (R, Python, SPSS) to tag each user with their group and record their scores for each software/dimension.
- Example Python code for tests:
from scipy.stats import mannwhitneyu, wilcoxon, friedmanchisquare # Mann-Whitney for Group A vs Group B on Dimension 1 u_stat, p_val = mannwhitneyu(groupA_soft1_dim1, groupB_soft2_dim1) print(f"Mann-Whitney U: {u_stat}, p-value: {p_val}") # Wilcoxon for Group C comparing Software 1 vs 2 on Dimension 1 w_stat, p_val = wilcoxon(groupC_soft1_dim1, groupC_soft2_dim1) print(f"Wilcoxon W: {w_stat}, p-value: {p_val}") # Friedman Test for Group C comparing all three software on Dimension 1 f_stat, p_val = friedmanchisquare(groupC_soft1_dim1, groupC_soft2_dim1, groupC_yoursoft_dim1) print(f"Friedman Chi-square: {f_stat}, p-value: {p_val}") - Visualization example (Python):
import seaborn as sns import matplotlib.pyplot as plt # Assuming your data is in a pandas DataFrame called 'test_data' sns.boxplot(x="software", y="dimension1", hue="user_group", data=test_data) plt.title("Dimension 1 Ratings by Software and User Group") plt.show() - Add qualitative context: If you can, follow up with a few open-ended questions (e.g., "What made you rate [Software X] higher on usability?")—this will help explain why you see statistical differences, which is crucial with a small sample.
内容的提问来源于stack exchange,提问作者user2721

