You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用sklearn NMF组件手动重构新数据与inverse_transform结果不一致

Why Your NMF Reconstruction Results Don't Match

Oh I see exactly where the mismatch is coming from—your manual calculation is doing a standard unconstrained linear projection, but scikit-learn's NMF relies on a non-negative constrained least-squares approach for its transform step, which changes everything. Let's break this down:

The Core Logic Gap

First, let's clarify how scikit-learn's NMF works under the hood:

  • After fitting, model.components_ is the basis matrix H (shape: n_components × n_features), which defines the latent components of the decomposition.
  • When you call model.transform(new_data), it solves a non-negative least-squares problem for each sample: it finds a non-negative weight matrix W_new (shape: n_samples × n_components) such that ||new_data - W_new @ H||² is minimized, with the hard constraint that all values in W_new are ≥ 0.
  • model.inverse_transform(W_new) then computes W_new @ H to produce the reconstructed data.

Your manual code, by contrast, calculates the orthogonal projection of new_data onto the space spanned by the NMF components—this ignores the non-negativity rule that's foundational to NMF. The unconstrained linear projection will often produce negative weights (which NMF explicitly forbids), leading to a completely different reconstruction.

How to Manually Reproduce Scikit-learn's Result

To get result_2 to match result_1, you need to replicate the non-negative least-squares step that scikit-learn uses for transform. You can do this easily with scipy.optimize.nnls for each sample:

import numpy as np
from scipy.optimize import nnls
from sklearn.decomposition import NMF
from sklearn.datasets import make_sparse_coded_signal

# Generate sample data for testing
X, _, _ = make_sparse_coded_signal(n_samples=10, n_components=5, n_features=20, random_state=42)
new_data = X[:3]

# Fit the NMF model
model = NMF(n_components=5, random_state=42)
model.fit(X)
H = model.components_

# Scikit-learn's built-in reconstruction
result_1 = model.inverse_transform(model.transform(new_data))

# Manual replication with non-negative least-squares
result_2 = []
for sample in new_data:
    # Solve min ||sample - w@H||² where w ≥ 0
    # nnls expects Ax = b, so we use H.T as A (since sample = w@H → sample = H.T @ w.T)
    weights, _ = nnls(H.T, sample)
    reconstructed_sample = weights @ H
    result_2.append(reconstructed_sample)
result_2 = np.array(result_2)

# Verify the match (should print True)
print(np.allclose(result_1, result_2))

Quick Notes

  • Scikit-learn's actual transform implementation uses more efficient algorithms (like coordinate descent or multiplicative updates) instead of per-sample nnls, but the core non-negative least-squares logic is identical.
  • Your original manual approach is great for linear subspace projection tasks, but it's fundamentally incompatible with NMF's non-negativity requirement—hence the mismatch you saw.

内容的提问来源于stack exchange,提问作者swathis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 10:29:33