You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python:拼接DataFrame后保留标识,拆分回原始数据集的方法问询

Solution for Splitting Processed Concatenated DataFrames Back to Originals

Hey, great question! Here's a straightforward and reliable approach to solve this problem—add a unique source identifier column to each original DataFrame before concatenation. This identifier will persist through your Geohash encoding process, making it trivial to split the processed X back into your original datasets later.

Step 1: Add Source Identifiers to Original DataFrames

First, assign a unique label to each of your X_1 to X_10 DataFrames. This label will act as your "split marker" throughout the workflow.

# Store your original DataFrames in a list for easy iteration
original_dfs = [X_1, X_2, X_3, X_4, X_5, X_6, X_7, X_8, X_9, X_10]

# Add a source identifier column to each DataFrame
for df_num, df in enumerate(original_dfs, start=1):
    # Use a string like "X_1" or just the number 1—choose what works best for you
    df['source_df'] = f"X_{df_num}"

# Concatenate into your combined X DataFrame
X = pd.concat(original_dfs, ignore_index=True)

Step 2: Run Your Geohash Encoding

Your existing encoding code will work exactly as before—since we're just adding a new column, the Geohash processing won't interfere with the source identifier:

Geo_as_Category = pd.Series(X['Geohash'], dtype="category")
X['GeohashFeature'] = Geo_as_Category.cat.codes

Step 3: Split Back to Original Datasets

Now you can easily filter the processed X using the source_df column to retrieve each original dataset (with the new GeohashFeature column included).

Option 1: Extract Individual DataFrames

# Retrieve each processed original DataFrame, removing the identifier column if needed
X_1_processed = X[X['source_df'] == "X_1"].drop('source_df', axis=1)
X_2_processed = X[X['source_df'] == "X_2"].drop('source_df', axis=1)
# ... repeat for X_3 to X_10

Option 2: Batch Extract into a Dictionary (More Efficient)

If you want to avoid repetitive code, store all processed DataFrames in a dictionary:

processed_dfs = {}
for df_num in range(1, 11):
    df_key = f"X_{df_num}_processed"
    # Filter and remove the identifier column (omit .drop() if you want to keep it)
    processed_dfs[df_key] = X[X['source_df'] == f"X_{df_num}"].drop('source_df', axis=1)

# Access processed datasets like this: processed_dfs['X_1_processed']

Bonus Optimization: Use Multi-Index Instead of a Column

If you don't want an extra column cluttering your DataFrame, you can set the source identifier as part of a multi-index instead:

# Add identifier to index during concatenation
X = pd.concat(
    [df for df in original_dfs],
    keys=[f"X_{i+1}" for i in range(10)],
    names=['source_df', 'row_idx']
)

# Split using xs() (cross-section)
X_1_processed = X.xs("X_1", level='source_df').reset_index(drop=True)

This keeps your main columns clean while still preserving the split marker.

内容的提问来源于stack exchange,提问作者Alex P

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 07:22:39