You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

尝试用FunctionTransformer为ColumnTransformer列加权触发ValueError

问题:为ColumnTransformer的列设置分类权重时触发维度错误

尝试为ColumnTransformer的列设置列间权重用于后续分类,由于无法直接在ColumnTransformer中实现该操作,采用numpy复制数值的方式实现,代码如下:

from numpy import repeat,ones
...
trf_list = [
  (
    feature,
    Pipeline([
      ('vectorizer', TfidfVectorizer(strip_accents='unicode', lowercase=True, stop_words='english', vocabulary=self.voc[feature])),
      ('weighter', FunctionTransformer(lambda row: repeat(row, weight)))
    ]),
    feature
  ) for feature, weight in zip(features, weights)]
trf = ColumnTransformer(transformers = trf_list, **kwargs)

test_data = np.array([[1, 9, 4], [2, 7, 1]])

expected_results = np.array([[1, 1, 9, 9, 4, 4], [2, 2, 7, 7, 1, 1]]) # 当weight=2时的预期结果

运行后触发ValueError,错误信息如下:

ValueError: The output of the 'description' transformer should be 2D (scipy matrix, array, or pandas DataFrame).
ValueError                                Traceback (most recent call last)
in engine
----> 1 m.fit(x_train)

/home/cdsw/src/V3/model.py in fit(self, train_set)
     65   
     66   def fit(self, train_set):
---> 67     self.pipeline.fit(train_set)
     68     
     69     

/home/cdsw/.local/lib/python3.6/site-packages/sklearn/pipeline.py in fit(self, X, y, **fit_params)
    339         """
    340         fit_params_steps = self._check_fit_params(**fit_params)
---> 341         Xt = self._fit(X, y, **fit_params_steps)
    342         with _print_elapsed_time('Pipeline',
    343                                  self._log_message(len(self.steps) - 1)):

/home/cdsw/.local/lib/python3.6/site-packages/sklearn/pipeline.py in _fit(self, X, y, **fit_params_steps)
    305                 message_clsname='Pipeline',
    306                 message=self._log_message(step_idx),
---> 307                 **fit_params_steps[name])
    308             # Replace the transformer of the step with the fitted
    309             # transformer. This is necessary when loading the transformer

/home/cdsw/.local/lib/python3.6/site-packages/joblib/memory.py in __call__(self, *args, **kwargs)
    347 
    348     def __call__(self, *args, **kwargs):
---> 349         return self.func(*args, **kwargs)
    350 
    351     def call_and_shelve(self, *args, **kwargs):

/home/cdsw/.local/lib/python3.6/site-packages/sklearn/pipeline.py in _fit_transform_one(transformer, X, y, weight, message_clsname, message, **fit_params)
    752     with _print_elapsed_time(message_clsname, message):
    753         if hasattr(transformer, 'fit_transform'):
---> 754             res = transformer.fit_transform(X, y, **fit_params)
    755         else:
    756             res = transformer.fit(X, y, **fit_params).transform(X)

/home/cdsw/.local/lib/python3.6/site-packages/sklearn/compose/_column_transformer.py in fit_transform(self, X, y)
    525 
    526         self._update_fitted_transformers(transformers)
---> 527         self._validate_output(Xs)
    528 
    529         return self._hstack(list(Xs))

/home/cdsw/.local/lib/python3.6/site-packages/sklearn/compose/_column_transformer.py in _validate_output(self, result)
    414                 raise ValueError(
    415                     "The output of the '{0}' transformer should be 2D (scipy "
---> 416                     "matrix, array, or pandas DataFrame).".format(name))
    417 
    418     def _log_message(self, name, idx, total):

ValueError: The output of the 'description' transformer should be 2D (scipy matrix, array, or pandas DataFrame).

(注:ColumnTransformer的输入为pandas dataframe)


解决方案

错误核心是FunctionTransformer中的repeat(row, weight)将原本的2D数组(形状为(n_samples, n_features))转换成了1D数组,而ColumnTransformer要求每个转换器的输出必须是2D结构。

要实现列复制并保持2D结构,只需修改repeat的使用方式,指定沿特征轴(axis=1)复制,同时保留样本维度:

  1. 修改FunctionTransformer的lambda函数:
('weighter', FunctionTransformer(lambda x: np.repeat(x, weight, axis=1)))
  1. 完整可运行代码示例:
import numpy as np
import pandas as pd
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.preprocessing import FunctionTransformer

# 模拟特征、权重与词汇表
features = ['col1', 'col2', 'col3']
weights = [2, 2, 2]
voc = {f: [f'word_{i}' for i in range(3)] for f in features}

trf_list = [
    (
        feature,
        Pipeline([
            ('vectorizer', TfidfVectorizer(strip_accents='unicode', lowercase=True, stop_words='english', vocabulary=voc[feature])),
            ('weighter', FunctionTransformer(lambda x: np.repeat(x, weight, axis=1)))
        ]),
        feature
    ) for feature, weight in zip(features, weights)]

trf = ColumnTransformer(transformers=trf_list)

# 测试数据
test_data = pd.DataFrame({
    'col1': ['text 1', 'text 2'],
    'col2': ['text 3', 'text 4'],
    'col3': ['text 5', 'text 6']
})

result = trf.fit_transform(test_data)
print(result.shape)  # 输出(2, 18),符合2D结构要求

修改后,每个转换器输出的数组将保持(n_samples, n_features*weight)的2D形状,既实现了按权重复制特征列的需求,也满足ColumnTransformer的维度要求。

内容的提问来源于stack exchange,提问作者Thomas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 09:24:27