如何在Python中计算Multicompare Tukey HSD?含嵌套列表数据需求
numpy.recarray Usage) Hey there! Let's walk through how to run Tukey's HSD (Honestly Significant Difference) test on your nested list data, and clear up the confusion around numpy.recarray—since it's not actually necessary for this task, but I'll explain why you might have thought it was.
Step 1: Prepare Your Data
First, let's formalize your sample nested list (I filled in a placeholder value for the fifth group to keep things consistent):
my_list_of_lists = [ [0.75, 0.78, 0.80, 0.77, 0.71, 0.69, 0.73], [0.76, 0.73, 0.88, 0.71, 0.72, 0.80, 0.72], [0.71, 0.75, 0.77, 0.79, 0.68, 0.77, 0.66], [0.72, 0.79, 0.82, 0.73, 0.75, 0.60, 0.72], [0.73, 0.71, 0.66, 0.79, 0.72, 0.67, 0.70] ]
We'll convert each sublist to a numpy array (scipy works with regular lists too, but arrays add consistency).
Step 2: Use Scipy's Tukey HSD Functions
The simplest way to run Tukey's multicompare test is with scipy.stats—there are two straightforward methods depending on how you prefer to structure your data.
Option 1: Direct Group Input with tukey_hsd
This function accepts each group of data as a separate argument, which pairs perfectly with your nested list:
import numpy as np from scipy.stats import tukey_hsd # Convert each sublist to a numpy array groups = [np.array(group) for group in my_list_of_lists] # Run Tukey's HSD test tukey_result = tukey_hsd(*groups) # Print the full results (pairwise comparisons, p-values, significance) print(tukey_result)
Option 2: Long-Format Data with pairwise_tukeyhsd
If you want to generate visualizations or work with a single flat dataset, use this method (it requires all values plus corresponding group labels):
from scipy.stats import pairwise_tukeyhsd # Combine all data into one flat array all_values = np.concatenate(my_list_of_lists) # Create group labels (e.g., "Group 1", "Group 2"...) group_labels = [] for group_num in range(len(my_list_of_lists)): group_labels.extend([f"Group {group_num+1}"] * len(my_list_of_lists[group_num])) group_labels = np.array(group_labels) # Run the test tukey_result = pairwise_tukeyhsd(all_values, group_labels) # Print results and plot simultaneous confidence intervals print(tukey_result) tukey_result.plot_simultaneous()
What About numpy.recarray?
You mentioned looking into numpy.recarray—here's the key context: recarrays are structured arrays that let you access columns by name (like a table), but they require equal-length columns. This is a problem if your sublists have different lengths (your fifth group had an ellipsis, suggesting it might be shorter than others).
Even if all groups were the same length, recarrays don't add any value for this Tukey test—scipy's functions work perfectly with regular numpy arrays or raw lists. If you're curious how to convert your data to a recarray (just for learning), here's how (but don't use this for the test itself):
# Only works if all sublists are the same length! dtype = [(f"group{i+1}", float) for i in range(len(my_list_of_lists))] rec_array = np.rec.fromarrays(my_list_of_lists, dtype=dtype)
Quick Assumption Check
Before trusting the results, make sure your data meets Tukey's assumptions:
- Normality within each group (use Shapiro-Wilk test to verify)
- Homogeneity of variances (use Bartlett's or Levene's test)
内容的提问来源于stack exchange,提问作者Steve Jade

