如何在PyTorch中计算加权平均及基于神经网络的加权平均?需编写C语言代码
Got it, let's break this down step by step. I'll cover basic weighted averaging in PyTorch, how to integrate it into a neural network, and a corresponding C implementation for low-level use cases.
1. 基础张量加权平均
For simple tensor operations, weighted averaging boils down to calculating the weighted sum of elements and dividing by the total weight. Here's how to do it for single and batch inputs:
单个张量的加权平均
import torch # 输入张量和对应的权重 input_tensor = torch.tensor([1.0, 2.0, 3.0, 4.0]) weights = torch.tensor([0.1, 0.2, 0.3, 0.4]) # 计算加权平均 weighted_sum = torch.sum(input_tensor * weights) total_weight = torch.sum(weights) weighted_avg = weighted_sum / total_weight print(f"加权平均结果: {weighted_avg.item()}") # 输出: 3.0
批量张量的加权平均
If you're working with batches (e.g., shape (batch_size, num_features)), compute the average along the feature dimension:
batch_input = torch.tensor([[1.0,2.0,3.0], [4.0,5.0,6.0]]) batch_weights = torch.tensor([[0.2,0.3,0.5], [0.1,0.4,0.5]]) # 沿特征维度(dim=1)计算加权和与总权重 weighted_sums = torch.sum(batch_input * batch_weights, dim=1) total_weights = torch.sum(batch_weights, dim=1) batch_weighted_avg = weighted_sums / total_weights print(f"批量加权平均结果: {batch_weighted_avg}") # 输出: tensor([2.3000, 5.3000])
Note: If your weights are already normalized (sum to 1), you can skip dividing by
total_weightand just usetorch.sum(input * weights).
2. 神经网络中的加权平均实现
Weighted averaging can be incorporated into neural networks in three common ways: fixed weights, learnable weights, or dynamically computed weights (like attention mechanisms).
2.1 固定权重的加权平均层
Create a reusable layer with static weights that don't update during training:
import torch.nn as nn class FixedWeightedAverage(nn.Module): def __init__(self, weights): super().__init__() # 注册为不可训练的参数 self.weights = nn.Parameter(torch.tensor(weights), requires_grad=False) def forward(self, x): # x形状: (batch_size, num_features),输出形状: (batch_size,) weighted_sum = torch.sum(x * self.weights, dim=1) return weighted_sum / torch.sum(self.weights) # 使用示例 layer = FixedWeightedAverage([0.1, 0.3, 0.6]) input_batch = torch.randn(5, 3) # 5个样本,每个3个特征 output = layer(input_batch) print(f"固定权重层输出形状: {output.shape}") # 输出: torch.Size([5])
2.2 可学习权重的加权平均层
Make weights trainable so the network can optimize them during training (use softmax to ensure weights sum to 1):
class LearnableWeightedAverage(nn.Module): def __init__(self, num_features): super().__init__() # 初始化权重,用正态分布 self.weights = nn.Parameter(torch.randn(num_features), requires_grad=True) def forward(self, x): # 归一化权重,避免数值不稳定 normalized_weights = torch.softmax(self.weights, dim=0) # 计算加权平均 return torch.sum(x * normalized_weights, dim=1) # 训练示例 layer = LearnableWeightedAverage(3) optimizer = torch.optim.SGD(layer.parameters(), lr=0.01) input_batch = torch.randn(5, 3) target = torch.randn(5) # 一次训练迭代 optimizer.zero_grad() output = layer(input_batch) loss = nn.MSELoss()(output, target) loss.backward() optimizer.step() print(f"训练后归一化权重: {torch.softmax(layer.weights, dim=0).detach().numpy()}")
2.3 动态权重的加权平均(注意力风格)
Compute weights dynamically for each input sample using a small sub-network (common in attention-based models):
class DynamicWeightedAverage(nn.Module): def __init__(self, num_features, hidden_dim=16): super().__init__() # 用MLP生成每个样本的权重 self.weight_generator = nn.Sequential( nn.Linear(num_features, hidden_dim), nn.ReLU(), nn.Linear(hidden_dim, num_features), nn.Softmax(dim=1) # 对每个样本的特征维度归一化 ) def forward(self, x): # x形状: (batch_size, num_features) weights = self.weight_generator(x) # 生成每个样本的权重,形状: (batch_size, num_features) weighted_avg = torch.sum(x * weights, dim=1) # 输出形状: (batch_size,) return weighted_avg # 使用示例 layer = DynamicWeightedAverage(3) input_batch = torch.randn(5, 3) output = layer(input_batch) print(f"动态权重层输出形状: {output.shape}") print(f"对应样本的权重:\n{layer.weight_generator(input_batch).detach().numpy()}")
For low-level applications, here's a C implementation that handles 1D arrays and batch (2D) arrays, with error handling for zero total weights:
#include <stdio.h> #include <stdlib.h> #include <math.h> // 一维数组的加权平均计算 float weighted_average_1d(const float* input, const float* weights, int length) { if (length <= 0) { printf("Error: Invalid array length\n"); return NAN; } float weighted_sum = 0.0f; float total_weight = 0.0f; for (int i = 0; i < length; i++) { weighted_sum += input[i] * weights[i]; total_weight += weights[i]; } // 处理权重总和为0的情况 if (fabs(total_weight) < 1e-6) { printf("Warning: Total weight is zero, returning NaN\n"); return NAN; } return weighted_sum / total_weight; } // 批量数组的加权平均计算,结果存储在result数组中 void weighted_average_batch(const float* input, const float* weights, int batch_size, int feature_size, float* result) { if (batch_size <= 0 || feature_size <= 0) { printf("Error: Invalid batch or feature size\n"); return; } for (int b = 0; b < batch_size; b++) { float ws = 0.0f; float tw = 0.0f; int start_idx = b * feature_size; for (int f = 0; f < feature_size; f++) { int idx = start_idx + f; ws += input[idx] * weights[idx]; tw += weights[idx]; } if (fabs(tw) < 1e-6) { result[b] = NAN; printf("Warning: Batch %d has zero total weight\n", b); } else { result[b] = ws / tw; } } } // 测试函数 int main() { // 一维数组测试 float input_1d[] = {1.0f, 2.0f, 3.0f, 4.0f}; float weights_1d[] = {0.1f, 0.2f, 0.3f, 0.4f}; float avg_1d = weighted_average_1d(input_1d, weights_1d, 4); printf("1D Weighted Average: %.4f\n", avg_1d); // 批量数组测试 int batch_size = 2; int feature_size = 3; float input_batch[] = {1.0f,2.0f,3.0f, 4.0f,5.0f,6.0f}; float weights_batch[] = {0.2f,0.3f,0.5f, 0.1f,0.4f,0.5f}; float* result_batch = (float*)malloc(batch_size * sizeof(float)); weighted_average_batch(input_batch, weights_batch, batch_size, feature_size, result_batch); printf("\nBatch Weighted Averages:\n"); for (int i = 0; i < batch_size; i++) { printf("Sample %d: %.4f\n", i, result_batch[i]); } free(result_batch); return 0; }
内容的提问来源于stack exchange,提问作者Driss AL

