Python嵌套循环仅保留最后轮结果:DataFrame聚类集标记异常
It looks like your issue stems from not creating an explicit copy of your DataFrame slice and potentially running into pandas' SettingWithCopy behavior, where changes to a view of the original DataFrame don't persist correctly. Additionally, initializing your new columns upfront can help avoid unexpected behavior when assigning values.
Here's the fixed approach:
Step 1: Create a copy of your DataFrame and initialize new columns
First, make sure you're working with a copy of the relevant columns (not just a view), and set up the cluster and cluster set columns with default values:
import pandas as pd # Create a copy of the relevant columns to avoid SettingWithCopy issues df_triplets = df[['Activity', 'SMILES']].copy() # Initialize new columns with nullable values df_triplets['cluster'] = pd.NA df_triplets['cluster set'] = pd.NA
Step 2: Refactor the loop for better performance and correctness
Instead of nesting three loops (which is slow for large datasets), you can use pandas' vectorized indexing to assign values to all molecules in a cluster at once. This is more efficient and avoids potential issues with individual row assignments:
list_points = [train_points, test_points, val_points] name_points = ['train', 'test', 'val'] for name, points in zip(name_points, list_points): for cluster_num, cluster_indices in enumerate(points): # Assign cluster number to all molecules in the cluster df_triplets.loc[cluster_indices, 'cluster'] = cluster_num # Assign cluster set to all molecules in the cluster df_triplets.loc[cluster_indices, 'cluster set'] = name
Why this works:
- Using
.copy()ensures you're modifying a separate DataFrame, not a view of the original one, so changes persist correctly across all loop iterations. - Initializing columns upfront means pandas doesn't have to dynamically add them during the loop, which can cause inconsistent behavior.
- Vectorized assignments (using a list of indices in
loc) are significantly faster than looping through each molecule individually, especially with large datasets.
Verify the results
You can confirm all sets are properly assigned with:
print(df_triplets['cluster set'].value_counts())
This should show counts for train, test, and val instead of just val.
内容的提问来源于stack exchange,提问作者Daniel

