You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

KMeans肘方法报错:输入含NaN/无穷大或超出float64范围

Troubleshooting ValueError in KMeans Elbow Method & Memory Error in Hierarchical Clustering

Let's work through your issues one by one, starting with the core ValueError blocking your KMeans elbow method:

1. Fix Your NaN/Infinity Check First

Your existing code for checking NaNs and infinite values has a logical flaw—np.isnan(df3.any()) and np.isfinite(df3.all()) don't actually scan your raw data for anomalies. Here's why:

  • df3.any() returns a boolean array indicating if each column has at least one True value
  • Wrapping that in np.isnan checks if those boolean values are NaN (which they never will be)

To properly scan your dataframe for NaNs or non-finite values (inf/-inf), use these lines instead:

# Check if any cell in the dataframe is NaN
print(np.isnan(df3).any().any())
# Check if all cells are finite (no NaN/inf/-inf)
print(np.isfinite(df3).all().all())

If either check returns False, clean your data with:

# Replace inf/-inf with NaN, then drop rows with any NaN
df_clean = df3.replace([np.inf, -np.inf], np.nan).dropna()
X = df_clean.values  # Use this cleaned array for KMeans

2. Resolve "Value Too Large" Issues (Even Without Inf)

Even if you don't have NaNs/inf, extreme float values can cause overflow during KMeans' inertia (WCSS) calculation. Log-transforming helps, but standardization/normalization is the better fix—it scales your data to a consistent range, eliminating overflow risks and improving clustering performance:

from sklearn.preprocessing import StandardScaler
# Scale data to mean=0, variance=1
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)

# Run elbow method with scaled data
wcss = []
for i in range(1, 11):
    kmeans = KMeans(n_clusters=i, init='k-means++', random_state=42)
    kmeans.fit(X_scaled)
    wcss.append(kmeans.inertia_)

# Your plotting code remains the same
plt.plot(range(1, 11), wcss)
plt.title('The Elbow Method')
plt.xlabel('Number of clusters')
plt.ylabel('WCSS')
plt.show()

If you have outliers, use RobustScaler instead—it's less sensitive to extreme values than StandardScaler.

3. Skip the Integer Conversion (It's Not the Solution)

Your attempts to convert float64 to int64 failed because your data has decimal values (e.g., 1.4494), which can't be safely cast to integers without losing information. Worse, this approach is unnecessary—KMeans works perfectly with float64 data, so abandon this line of troubleshooting entirely.


4. Fix the Hierarchical Clustering Memory Error

The MemoryError: Unable to allocate 722. GiB happens because hierarchical clustering with the ward method requires computing a full pairwise distance matrix, which has a space complexity of O(n²). For large datasets (yours looks to have ~440k samples), this is computationally impossible. Here are workarounds:

  • Sample your data: Draw a random subset of 5k-10k samples to generate the dendrogram, identify a suitable cluster count, then apply that count to your full dataset using KMeans or MiniBatchKMeans.
  • Use dimensionality reduction: Run PCA first to reduce your features to 2-3 dimensions, then perform hierarchical clustering on the reduced data. This drastically cuts down the distance matrix size.
  • Avoid full dendrograms: If you don't need the visual, use scipy.cluster.hierarchy.fclusterdata to directly assign clusters without generating the full dendrogram.

Step-by-Step Validation

  1. Run the corrected NaN/inf checks and clean your data if needed.
  2. Scale your data and re-run the KMeans elbow method.
  3. For hierarchical clustering, use sampling or dimensionality reduction to avoid memory overload.

内容的提问来源于stack exchange,提问作者Killi Mandjaro

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 18:22:47