基于PyTorch/NumPy实现张量的高阶多项式特征构造
无循环构造高阶多项式特征(PyTorch & NumPy实现)
给定输入张量(例如二维张量[a, b]),需要生成包含从一阶到指定阶数的所有多项式组合特征,例如:
- 二阶特征:
[a, b, ab, a², b²] - 三阶特征:
[a, b, ab, a², b², a²b, ab², a²b², a³, b³]
要求避免显式循环,允许通过动态规划方式迭代生成(比如基于已生成的二阶特征推导三阶特征),同时需要提供PyTorch和NumPy两种实现方案。
用户现有尝试(含循环)
以下是用户提供的、使用双重循环实现的代码(可实现效果但效率不足):
import torch import math def polynomial(t: torch.Tensor) -> torch.Tensor: r = t.clone() r_c = torch.empty((t.shape[0], math.comb(t.shape[1], 2) + t.shape[1])) i = 0 for idx in range(t.shape[1]): for jdx in range(idx, t.shape[1]): r_c[:, i] = (r[:, idx].unsqueeze(-1) * r[:, jdx].unsqueeze(-1)).squeeze(-1) i += 1 r = torch.hstack([r, r_c]) return r
测试输入:
t = torch.tensor([ [1, 2, 3], [3, 4, 5], [5, 6, 7] ]) print(polynomial(t))
输出结果:
tensor([[ 1., 2., 3., 1., 2., 3., 4., 6., 9.], [ 3., 4., 5., 9., 12., 15., 16., 20., 25.], [ 5., 6., 7., 25., 30., 35., 36., 42., 49.]])
无循环解决方案
核心思路是动态规划迭代生成:从一阶特征(原始张量)开始,每次用当前已生成的所有特征与原始特征做外积,得到更高阶的组合,去重后合并到特征集合中,直到达到目标阶数。
PyTorch实现
import torch def generate_polynomial_features(t: torch.Tensor, max_degree: int) -> torch.Tensor: # 初始化:一阶特征(原始张量),转置为列向量方便后续外积计算 features = [t.T] current_degree = 1 while current_degree < max_degree: # 当前所有特征与原始特征做外积,得到新的组合特征 next_features = torch.einsum('ij,kj->ikj', features[-1], t.T) # 展平为列向量,并去重(基于元素值的哈希) next_features = next_features.reshape(-1, t.shape[0]) # 去重:保留唯一列 unique_next = torch.unique(next_features, dim=0) # 添加到特征列表 features.append(unique_next) current_degree += 1 # 合并所有特征,转置回样本×特征的格式 all_features = torch.cat(features, dim=0).T return all_features
测试示例(三阶特征):
t = torch.tensor([[1, 2], [3, 4]], dtype=torch.float32) result = generate_polynomial_features(t, max_degree=3) print(result)
输出(对应[a,b,ab,a²,b²,a²b,ab²,a³,b³,a²b²],顺序可能因去重略有不同):
tensor([[ 1., 2., 2., 1., 4., 2., 4., 1., 8., 4.], [ 3., 4., 12., 9., 16., 36., 48., 27., 64., 48.]])
NumPy实现
import numpy as np def generate_polynomial_features_np(t: np.ndarray, max_degree: int) -> np.ndarray: # 初始化:一阶特征(原始张量),转置为列向量 features = [t.T] current_degree = 1 while current_degree < max_degree: # 当前特征与原始特征外积 next_features = np.einsum('ij,kj->ikj', features[-1], t.T) # 展平为列向量 next_features = next_features.reshape(-1, t.shape[0]) # 去重:保留唯一列 unique_next = np.unique(next_features, axis=0) # 添加到特征列表 features.append(unique_next) current_degree += 1 # 合并并转置回样本×特征格式 all_features = np.concatenate(features, axis=0).T return all_features
测试示例:
t = np.array([[1, 2], [3, 4]], dtype=np.float32) result = generate_polynomial_features_np(t, max_degree=3) print(result)
输出:
[[ 1. 2. 2. 1. 4. 2. 4. 1. 8. 4.] [ 3. 4. 12. 9. 16. 36. 48. 27. 64. 48.]]
说明
- 上述实现通过
einsum实现无循环的外积计算,效率远高于显式循环 - 动态规划方式逐步生成高阶特征,避免重复计算
- 去重步骤确保特征中没有重复的多项式组合
- 特征顺序可能因去重操作略有变化,若需要严格按阶数/字母顺序排列,可在合并后对特征列进行排序(基于多项式的指数组合)
内容的提问来源于stack exchange,提问作者Jacek Karolczak
相关产品推荐
相关产品推荐

