You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python嵌套循环仅保留最后轮结果:DataFrame聚类集标记异常

It looks like your issue stems from not creating an explicit copy of your DataFrame slice and potentially running into pandas' SettingWithCopy behavior, where changes to a view of the original DataFrame don't persist correctly. Additionally, initializing your new columns upfront can help avoid unexpected behavior when assigning values.

Here's the fixed approach:

Step 1: Create a copy of your DataFrame and initialize new columns

First, make sure you're working with a copy of the relevant columns (not just a view), and set up the cluster and cluster set columns with default values:

import pandas as pd

# Create a copy of the relevant columns to avoid SettingWithCopy issues
df_triplets = df[['Activity', 'SMILES']].copy()

# Initialize new columns with nullable values
df_triplets['cluster'] = pd.NA
df_triplets['cluster set'] = pd.NA

Step 2: Refactor the loop for better performance and correctness

Instead of nesting three loops (which is slow for large datasets), you can use pandas' vectorized indexing to assign values to all molecules in a cluster at once. This is more efficient and avoids potential issues with individual row assignments:

list_points = [train_points, test_points, val_points]
name_points = ['train', 'test', 'val']

for name, points in zip(name_points, list_points):
    for cluster_num, cluster_indices in enumerate(points):
        # Assign cluster number to all molecules in the cluster
        df_triplets.loc[cluster_indices, 'cluster'] = cluster_num
        # Assign cluster set to all molecules in the cluster
        df_triplets.loc[cluster_indices, 'cluster set'] = name

Why this works:

  • Using .copy() ensures you're modifying a separate DataFrame, not a view of the original one, so changes persist correctly across all loop iterations.
  • Initializing columns upfront means pandas doesn't have to dynamically add them during the loop, which can cause inconsistent behavior.
  • Vectorized assignments (using a list of indices in loc) are significantly faster than looping through each molecule individually, especially with large datasets.

Verify the results

You can confirm all sets are properly assigned with:

print(df_triplets['cluster set'].value_counts())

This should show counts for train, test, and val instead of just val.

内容的提问来源于stack exchange,提问作者Daniel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:49:40