Python:拼接DataFrame后保留标识,拆分回原始数据集的方法问询
Hey, great question! Here's a straightforward and reliable approach to solve this problem—add a unique source identifier column to each original DataFrame before concatenation. This identifier will persist through your Geohash encoding process, making it trivial to split the processed X back into your original datasets later.
Step 1: Add Source Identifiers to Original DataFrames
First, assign a unique label to each of your X_1 to X_10 DataFrames. This label will act as your "split marker" throughout the workflow.
# Store your original DataFrames in a list for easy iteration original_dfs = [X_1, X_2, X_3, X_4, X_5, X_6, X_7, X_8, X_9, X_10] # Add a source identifier column to each DataFrame for df_num, df in enumerate(original_dfs, start=1): # Use a string like "X_1" or just the number 1—choose what works best for you df['source_df'] = f"X_{df_num}" # Concatenate into your combined X DataFrame X = pd.concat(original_dfs, ignore_index=True)
Step 2: Run Your Geohash Encoding
Your existing encoding code will work exactly as before—since we're just adding a new column, the Geohash processing won't interfere with the source identifier:
Geo_as_Category = pd.Series(X['Geohash'], dtype="category") X['GeohashFeature'] = Geo_as_Category.cat.codes
Step 3: Split Back to Original Datasets
Now you can easily filter the processed X using the source_df column to retrieve each original dataset (with the new GeohashFeature column included).
Option 1: Extract Individual DataFrames
# Retrieve each processed original DataFrame, removing the identifier column if needed X_1_processed = X[X['source_df'] == "X_1"].drop('source_df', axis=1) X_2_processed = X[X['source_df'] == "X_2"].drop('source_df', axis=1) # ... repeat for X_3 to X_10
Option 2: Batch Extract into a Dictionary (More Efficient)
If you want to avoid repetitive code, store all processed DataFrames in a dictionary:
processed_dfs = {} for df_num in range(1, 11): df_key = f"X_{df_num}_processed" # Filter and remove the identifier column (omit .drop() if you want to keep it) processed_dfs[df_key] = X[X['source_df'] == f"X_{df_num}"].drop('source_df', axis=1) # Access processed datasets like this: processed_dfs['X_1_processed']
Bonus Optimization: Use Multi-Index Instead of a Column
If you don't want an extra column cluttering your DataFrame, you can set the source identifier as part of a multi-index instead:
# Add identifier to index during concatenation X = pd.concat( [df for df in original_dfs], keys=[f"X_{i+1}" for i in range(10)], names=['source_df', 'row_idx'] ) # Split using xs() (cross-section) X_1_processed = X.xs("X_1", level='source_df').reset_index(drop=True)
This keeps your main columns clean while still preserving the split marker.
内容的提问来源于stack exchange,提问作者Alex P

