R语言:基于两个列表计算各月份的加权均值
Alright, let's figure out how to calculate weighted means for all your months in bulk. Based on your example for April2002, the core idea is multiplying each data point by its normalized weight (weight divided by total weight) and summing those up. Here are a couple of straightforward ways to do this for all your months at once:
Method 1: Pure Python (No External Libraries)
First, let's structure your data clearly. We'll use a dictionary to map each month to its corresponding data points, and a separate list for weights.
Step 1: Define Your Data
# Dictionary mapping month names to their data points monthly_data = { "April2002": [0.05, 0.01, 0.02, 0.1], "April2003": [0.03, 0.04, 0.06, 0.08], "May2002": [0.07, 0.02, 0.03, 0.09], # Add as many months as you need here } # Your weight list (matches the length of each month's data points) weights = [10, 20, 25, 15]
Step 2: Calculate Total Weight
We'll precompute the total weight once to avoid redundant calculations:
total_weight = sum(weights)
Step 3: Define a Helper Function for Weighted Mean
def get_weighted_mean(data_points, weights, total_weight): # Calculate the sum of (normalized weight * data point) for all pairs return sum((weight / total_weight) * point for weight, point in zip(weights, data_points))
Step 4: Batch Process All Months
Loop through each month, compute its weighted mean, and store the results:
weighted_mean_results = {} for month, data in monthly_data.items(): # Quick check to make sure data and weights have the same length if len(data) != len(weights): print(f"⚠️ Skipping {month}: Data length doesn't match weight length") continue weighted_mean_results[month] = get_weighted_mean(data, weights, total_weight) # Print the final results for month, mean in weighted_mean_results.items(): print(f"Weighted mean for {month}: {round(mean, 4)}")
Method 2: Using Pandas (For Larger Datasets)
If you're working with a lot of months or data points, Pandas makes this even cleaner and faster:
Step 1: Set Up the DataFrame
import pandas as pd # Create a DataFrame where each column is a month's data df = pd.DataFrame({ "April2002": [0.05, 0.01, 0.02, 0.1], "April2003": [0.03, 0.04, 0.06, 0.08], "May2002": [0.07, 0.02, 0.03, 0.09] }) weights = [10, 20, 25, 15] total_weight = sum(weights)
Step 2: Compute Weighted Means for All Columns
Use apply() to calculate the weighted mean for each month (column) in one line:
weighted_means = df.apply(lambda column: sum((weights[i]/total_weight)*column[i] for i in range(len(column)))) # Print the results print(weighted_means)
Both methods will give you the exact calculation you showed for April2002 (which works out to ~0.0386, by the way) and scale seamlessly to all your other months.
内容的提问来源于stack exchange,提问作者Bit

