尝试用FunctionTransformer为ColumnTransformer列加权触发ValueError
问题:为ColumnTransformer的列设置分类权重时触发维度错误
尝试为ColumnTransformer的列设置列间权重用于后续分类,由于无法直接在ColumnTransformer中实现该操作,采用numpy复制数值的方式实现,代码如下:
from numpy import repeat,ones ... trf_list = [ ( feature, Pipeline([ ('vectorizer', TfidfVectorizer(strip_accents='unicode', lowercase=True, stop_words='english', vocabulary=self.voc[feature])), ('weighter', FunctionTransformer(lambda row: repeat(row, weight))) ]), feature ) for feature, weight in zip(features, weights)] trf = ColumnTransformer(transformers = trf_list, **kwargs) test_data = np.array([[1, 9, 4], [2, 7, 1]]) expected_results = np.array([[1, 1, 9, 9, 4, 4], [2, 2, 7, 7, 1, 1]]) # 当weight=2时的预期结果
运行后触发ValueError,错误信息如下:
ValueError: The output of the 'description' transformer should be 2D (scipy matrix, array, or pandas DataFrame). ValueError Traceback (most recent call last) in engine ----> 1 m.fit(x_train) /home/cdsw/src/V3/model.py in fit(self, train_set) 65 66 def fit(self, train_set): ---> 67 self.pipeline.fit(train_set) 68 69 /home/cdsw/.local/lib/python3.6/site-packages/sklearn/pipeline.py in fit(self, X, y, **fit_params) 339 """ 340 fit_params_steps = self._check_fit_params(**fit_params) ---> 341 Xt = self._fit(X, y, **fit_params_steps) 342 with _print_elapsed_time('Pipeline', 343 self._log_message(len(self.steps) - 1)): /home/cdsw/.local/lib/python3.6/site-packages/sklearn/pipeline.py in _fit(self, X, y, **fit_params_steps) 305 message_clsname='Pipeline', 306 message=self._log_message(step_idx), ---> 307 **fit_params_steps[name]) 308 # Replace the transformer of the step with the fitted 309 # transformer. This is necessary when loading the transformer /home/cdsw/.local/lib/python3.6/site-packages/joblib/memory.py in __call__(self, *args, **kwargs) 347 348 def __call__(self, *args, **kwargs): ---> 349 return self.func(*args, **kwargs) 350 351 def call_and_shelve(self, *args, **kwargs): /home/cdsw/.local/lib/python3.6/site-packages/sklearn/pipeline.py in _fit_transform_one(transformer, X, y, weight, message_clsname, message, **fit_params) 752 with _print_elapsed_time(message_clsname, message): 753 if hasattr(transformer, 'fit_transform'): ---> 754 res = transformer.fit_transform(X, y, **fit_params) 755 else: 756 res = transformer.fit(X, y, **fit_params).transform(X) /home/cdsw/.local/lib/python3.6/site-packages/sklearn/compose/_column_transformer.py in fit_transform(self, X, y) 525 526 self._update_fitted_transformers(transformers) ---> 527 self._validate_output(Xs) 528 529 return self._hstack(list(Xs)) /home/cdsw/.local/lib/python3.6/site-packages/sklearn/compose/_column_transformer.py in _validate_output(self, result) 414 raise ValueError( 415 "The output of the '{0}' transformer should be 2D (scipy " ---> 416 "matrix, array, or pandas DataFrame).".format(name)) 417 418 def _log_message(self, name, idx, total): ValueError: The output of the 'description' transformer should be 2D (scipy matrix, array, or pandas DataFrame).
(注:ColumnTransformer的输入为pandas dataframe)
解决方案
错误核心是FunctionTransformer中的repeat(row, weight)将原本的2D数组(形状为(n_samples, n_features))转换成了1D数组,而ColumnTransformer要求每个转换器的输出必须是2D结构。
要实现列复制并保持2D结构,只需修改repeat的使用方式,指定沿特征轴(axis=1)复制,同时保留样本维度:
- 修改FunctionTransformer的lambda函数:
('weighter', FunctionTransformer(lambda x: np.repeat(x, weight, axis=1)))
- 完整可运行代码示例:
import numpy as np import pandas as pd from sklearn.compose import ColumnTransformer from sklearn.pipeline import Pipeline from sklearn.feature_extraction.text import TfidfVectorizer from sklearn.preprocessing import FunctionTransformer # 模拟特征、权重与词汇表 features = ['col1', 'col2', 'col3'] weights = [2, 2, 2] voc = {f: [f'word_{i}' for i in range(3)] for f in features} trf_list = [ ( feature, Pipeline([ ('vectorizer', TfidfVectorizer(strip_accents='unicode', lowercase=True, stop_words='english', vocabulary=voc[feature])), ('weighter', FunctionTransformer(lambda x: np.repeat(x, weight, axis=1))) ]), feature ) for feature, weight in zip(features, weights)] trf = ColumnTransformer(transformers=trf_list) # 测试数据 test_data = pd.DataFrame({ 'col1': ['text 1', 'text 2'], 'col2': ['text 3', 'text 4'], 'col3': ['text 5', 'text 6'] }) result = trf.fit_transform(test_data) print(result.shape) # 输出(2, 18),符合2D结构要求
修改后,每个转换器输出的数组将保持(n_samples, n_features*weight)的2D形状,既实现了按权重复制特征列的需求,也满足ColumnTransformer的维度要求。
内容的提问来源于stack exchange,提问作者Thomas
相关产品推荐
相关产品推荐

