如何实现基于用户输入指定样本的双样本t检验?
Custom User-Input Samples for Two-Sample T-Test
Hey there! Let's fix up your code so you can use user-specified sample values instead of generating random Gaussian distributions. Your initial input logic was targeting distribution parameters (like mean/variance), but we need to capture actual sample data instead. Here's a complete, working solution:
Full Modified Code
import numpy as np from scipy import stats # Get user input for Sample A sample_a_str = input("Enter Sample A values separated by commas (e.g., 1.2, 3.4, 5.6): ") # Convert input string to a numpy array of floats a = np.array([float(val.strip()) for val in sample_a_str.split(',')]) # Get user input for Sample B sample_b_str = input("Enter Sample B values separated by commas (e.g., 2.3, 4.5, 6.7): ") b = np.array([float(val.strip()) for val in sample_b_str.split(',')]) # Confirm the input samples print("\nSample A:", a) print("Sample B:", b) # Calculate sample sizes n_a = len(a) n_b = len(b) # Compute unbiased variances (ddof=1 uses n-1 for degrees of freedom) var_a = a.var(ddof=1) var_b = b.var(ddof=1) # Calculate pooled standard deviation (adjust for equal/unequal sample sizes) if n_a == n_b: pooled_std = np.sqrt((var_a + var_b) / 2) else: pooled_std = np.sqrt(((n_a - 1)*var_a + (n_b - 1)*var_b) / (n_a + n_b - 2)) # Compute t-statistic t_stat = (a.mean() - b.mean()) / (pooled_std * np.sqrt(1/n_a + 1/n_b)) # Calculate degrees of freedom df = n_a + n_b - 2 # Compute two-tailed p-value p_val = 2 * (1 - stats.t.cdf(np.abs(t_stat), df=df)) # Print manual calculation results print("\nManual Calculation Results:") print(f"t-statistic = {t_stat:.4f}") print(f"p-value = {p_val:.4f}") # Cross-validate with scipy's built-in function t_scipy, p_scipy = stats.ttest_ind(a, b, equal_var=True) print("\nScipy Built-in Function Results:") print(f"t-statistic = {t_scipy:.4f}") print(f"p-value = {p_scipy:.4f}")
Key Improvements & Explanations
- Sample Input Handling: We now accept comma-separated numerical values from the user, convert them into numpy arrays—this lets us directly work with the exact sample data you need.
- Flexible Sample Sizes: The code checks if sample sizes are equal and adjusts the pooled standard deviation formula accordingly, making it work for both equal and unequal sample sizes.
- Two-Tailed Test Alignment: The manual p-value calculation uses the absolute t-statistic and multiplies by 2, matching the default two-tailed output of
scipy.stats.ttest_ind. - Validation: We keep the scipy cross-check to ensure our manual calculations match the trusted library implementation.
Example Usage
If you input:
- Sample A:
2.1, 1.9, 2.3, 2.0, 1.8 - Sample B:
0.2, -0.1, 0.3, 0.0, -0.2
You'll get output showing a significant difference between the two sample means (low p-value), just like your original random sample example.
内容的提问来源于stack exchange,提问作者Petkes Balás
相关产品推荐
相关产品推荐

