使用NMF分解混合光谱时fit_transform报ValueError的解决请求
NMF分解混合光谱的ValueError修复方案
问题场景
用非负矩阵分解(NMF)拆分混合化合物光谱,mixed_spectrum为混合光谱的numpy数组,W_init包含已知纯化合物光谱列向量及未知成分的初始值,运行代码时触发维度错误。
原代码
import numpy as np from sklearn.decomposition import NMF pure_compounds /= np.max(pure_compounds) n_components = len(pure_compounds) + 1 model = NMF(n_components=n_components, init='custom') W_init = np.concatenate((pure_compounds, np.ones((pure_compounds.shape[0], 1))), axis=1) print(W_init) print(np.shape(W_init)) mix = mixed_spectrum.reshape(-1, 1) print(mix) print(np.shape(mix)) W = model.fit_transform(mix, W=W_init) H = model.components_ for i in range(n_components): coeff = W[:, i] / np.sum(W, axis=1) print('Concentration coefficients for compound', i, ':', coeff)
错误输出
[[0.00000000e+00 2.76922293e-04 1.44273203e-03 0.00000000e+00 1.00000000e+00] [0.00000000e+00 2.85985939e-04 1.40933148e-03 0.00000000e+00 1.00000000e+00] [0.00000000e+00 2.54579755e-04 1.52763958e-03 0.00000000e+00 1.00000000e+00] ... [1.00000000e-06 1.00000000e-06 1.00000000e-06 1.00000000e-06 1.00000000e+00] [1.00000000e-06 1.00000000e-06 1.00000000e-06 1.00000000e-06 1.00000000e+00] [1.00000000e-06 1.00000000e-06 1.00000000e-06 1.00000000e+00]] (8954, 5) [[2.31330983e-02] [1.00000000e-06] [1.61938395e-02] ... [2.42525353e-02] [1.95440454e-02] [2.98149646e-03]] (8954, 1) --------------------------------------------------------------------------- ValueError Traceback (most recent call last) ~\AppData\Local\Temp\ipykernel_6528\919808671.py in <module> 8 print(mix) 9 print(np.shape(mix)) ---> 10 W = model.fit_transform(mix, W=W_init) 11 H = model.components_ 12 for i in range(n_components): ~\Anaconda3\lib\site-packages\sklearn\decomposition\_nmf.py in fit_transform(self, X, y, W, H) 1536 1537 with config_context(assume_finite=True): -> 1538 W, H, n_iter = self._fit_transform(X, W=W, H=H) 1539 1540 self.reconstruction_err_ = _beta_divergence( ~\Anaconda3\lib\site-packages\sklearn\decomposition\_nmf.py in _fit_transform(self, X, y, W, H, update_H) 1595 1596 # initialize or check W and H -> 1597 W, H = self._check_w_h(X, W, H, update_H) 1598 1599 # scale the regularization terms ~\Anaconda3\lib\site-packages\sklearn\decomposition\_nmf.py in _check_w_h(self, X, W, H, update_H) 1460 n_samples, n_features = X.shape 1461 if self.init == "custom" and update_H: -> 1462 _check_init(H, (self._n_components, n_features), "NMF (input H)") 1463 _check_init(W, (n_samples, self._n_components), "NMF (input W)") 1464 if H.dtype != X.dtype or W.dtype != X.dtype: ~\Anaconda3\lib\site-packages\sklearn\decomposition\_nmf.py in _check_init(A, shape, whom) 52 53 def _check_init(A, shape, whom): -> 54 A = check_array(A) 55 if np.shape(A) != shape: 56 raise ValueError( ~\Anaconda3\lib\site-packages\sklearn\utils\validation.py in check_array(array, accept_sparse, accept_large_sparse, dtype, order, copy, force_all_finite, ensure_2d, allow_nd, ensure_min_samples, ensure_min_features, estimator) 759 # If input is scalar raise error 760 if array.ndim == 0: -> 761 raise ValueError( 762 "Expected 2D array, got scalar array instead:\narray={}.\n" 763 "Reshape your data either using array.reshape(-1, 1) if " ValueError: Expected 2D array, got scalar array instead: array=None. Reshape your data either using array.reshape(-1, 1) if your data has a single feature or array.reshape(1, -1) if it contains a single sample.
修复方案
报错核心原因:当NMF设置init='custom'时,scikit-learn要求同时传入W和H的初始值,仅传W_init会导致H默认值为None,触发维度校验错误。
修改步骤:
- 初始化H的初始矩阵,维度需匹配
(n_components, n_features),其中n_features是输入X的特征数(此处mix的shape为(8954,1),故n_features=1),用极小非负值初始化避免零值问题。 - 在
fit_transform中同时传入W和H的初始值。
修改后的代码:
import numpy as np from sklearn.decomposition import NMF pure_compounds /= np.max(pure_compounds) n_components = len(pure_compounds) + 1 model = NMF(n_components=n_components, init='custom') W_init = np.concatenate((pure_compounds, np.ones((pure_compounds.shape[0], 1))), axis=1) # 初始化H的初始矩阵,维度与模型要求一致 H_init = np.ones((n_components, 1)) * 1e-6 print(W_init.shape) mix = mixed_spectrum.reshape(-1, 1) print(mix.shape) # 同时传入W和H的初始值 W = model.fit_transform(mix, W=W_init, H=H_init) H = model.components_ for i in range(n_components): coeff = W[:, i] / np.sum(W, axis=1) print('Concentration coefficients for compound', i, ':', coeff)
额外注意:
- 你的
W_init维度为(8954,5),与n_components=5匹配,这部分无需调整。 - 若无需更新H,可设置
update_H=False,但该场景仅适用于已知H真实值的情况,不符合当前需求,因此初始化H并传入是更合理的选择。
内容的提问来源于stack exchange,提问作者gionti
相关产品推荐
相关产品推荐

