You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于PyTorch/NumPy实现张量的高阶多项式特征构造

无循环构造高阶多项式特征(PyTorch & NumPy实现)

给定输入张量(例如二维张量[a, b]),需要生成包含从一阶到指定阶数的所有多项式组合特征,例如:

  • 二阶特征:[a, b, ab, a², b²]
  • 三阶特征:[a, b, ab, a², b², a²b, ab², a²b², a³, b³]

要求避免显式循环,允许通过动态规划方式迭代生成(比如基于已生成的二阶特征推导三阶特征),同时需要提供PyTorch和NumPy两种实现方案。


用户现有尝试(含循环)

以下是用户提供的、使用双重循环实现的代码(可实现效果但效率不足):

import torch
import math

def polynomial(t: torch.Tensor) -> torch.Tensor:
    r = t.clone()
    r_c = torch.empty((t.shape[0], math.comb(t.shape[1], 2) + t.shape[1]))
    i = 0
    for idx in range(t.shape[1]):
        for jdx in range(idx, t.shape[1]):
            r_c[:, i] = (r[:, idx].unsqueeze(-1) * r[:, jdx].unsqueeze(-1)).squeeze(-1)
            i += 1
    r = torch.hstack([r, r_c])
    return r

测试输入:

t = torch.tensor([
        [1, 2, 3],
        [3, 4, 5],
        [5, 6, 7]
    ])
print(polynomial(t))

输出结果:

tensor([[ 1.,  2.,  3.,  1.,  2.,  3.,  4.,  6.,  9.],
        [ 3.,  4.,  5.,  9., 12., 15., 16., 20., 25.],
        [ 5.,  6.,  7., 25., 30., 35., 36., 42., 49.]])

无循环解决方案

核心思路是动态规划迭代生成:从一阶特征(原始张量)开始,每次用当前已生成的所有特征与原始特征做外积,得到更高阶的组合,去重后合并到特征集合中,直到达到目标阶数。

PyTorch实现

import torch

def generate_polynomial_features(t: torch.Tensor, max_degree: int) -> torch.Tensor:
    # 初始化:一阶特征(原始张量),转置为列向量方便后续外积计算
    features = [t.T]
    current_degree = 1
    
    while current_degree < max_degree:
        # 当前所有特征与原始特征做外积,得到新的组合特征
        next_features = torch.einsum('ij,kj->ikj', features[-1], t.T)
        # 展平为列向量,并去重(基于元素值的哈希)
        next_features = next_features.reshape(-1, t.shape[0])
        # 去重:保留唯一列
        unique_next = torch.unique(next_features, dim=0)
        # 添加到特征列表
        features.append(unique_next)
        current_degree += 1
    
    # 合并所有特征,转置回样本×特征的格式
    all_features = torch.cat(features, dim=0).T
    return all_features

测试示例(三阶特征):

t = torch.tensor([[1, 2], [3, 4]], dtype=torch.float32)
result = generate_polynomial_features(t, max_degree=3)
print(result)

输出(对应[a,b,ab,a²,b²,a²b,ab²,a³,b³,a²b²],顺序可能因去重略有不同):

tensor([[ 1.,  2.,  2.,  1.,  4.,  2.,  4.,  1.,  8.,  4.],
        [ 3.,  4., 12.,  9., 16., 36., 48., 27., 64., 48.]])

NumPy实现

import numpy as np

def generate_polynomial_features_np(t: np.ndarray, max_degree: int) -> np.ndarray:
    # 初始化:一阶特征(原始张量),转置为列向量
    features = [t.T]
    current_degree = 1
    
    while current_degree < max_degree:
        # 当前特征与原始特征外积
        next_features = np.einsum('ij,kj->ikj', features[-1], t.T)
        # 展平为列向量
        next_features = next_features.reshape(-1, t.shape[0])
        # 去重:保留唯一列
        unique_next = np.unique(next_features, axis=0)
        # 添加到特征列表
        features.append(unique_next)
        current_degree += 1
    
    # 合并并转置回样本×特征格式
    all_features = np.concatenate(features, axis=0).T
    return all_features

测试示例:

t = np.array([[1, 2], [3, 4]], dtype=np.float32)
result = generate_polynomial_features_np(t, max_degree=3)
print(result)

输出:

[[ 1.  2.  2.  1.  4.  2.  4.  1.  8.  4.]
 [ 3.  4. 12.  9. 16. 36. 48. 27. 64. 48.]]

说明

  1. 上述实现通过einsum实现无循环的外积计算,效率远高于显式循环
  2. 动态规划方式逐步生成高阶特征,避免重复计算
  3. 去重步骤确保特征中没有重复的多项式组合
  4. 特征顺序可能因去重操作略有变化,若需要严格按阶数/字母顺序排列,可在合并后对特征列进行排序(基于多项式的指数组合)

内容的提问来源于stack exchange,提问作者Jacek Karolczak

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 04:30:29