如何将数据集A列值匹配对应Emb_A结果并新增列?
Got it, let's walk through how to add the Emb_A column to your dataset by mapping values from column A to their corresponding embeddings. Here's a straightforward solution using pandas, which is perfect for this kind of data matching task:
Step 1: Set up your data and mapping
First, we'll load your original dataset and create a dictionary for the A-to-Emb_A mapping. This dictionary makes lookups fast and easy.
import pandas as pd # Load your original dataset original_df = pd.DataFrame({ 'id': [1, 1, 1, 1, 1, 1, 1, 1, 1], 'tslot': [ '2014-11-02 22:45:00', '2014-10-26 09:15:00', '2014-10-26 11:00:00', '2014-10-26 11:15:00', '2014-10-26 12:00:00', '2014-10-26 13:00:00', '2014-10-26 13:45:00', '2014-10-26 14:45:00', '2014-10-26 15:45:00' ], 'A': [89, 762, 25, 762, 25, 25, 25, 902, 902] }) # Create the A-to-Emb_A mapping dictionary embedding_map = { 762.0: -0.237305, 89.0: -0.033929, 902.0: -0.237330, 25.0: -0.237300, 1176.0: -0.237300, 827.0: -0.237312, 1227.0: -0.237301 }
Step 2: Match data types and add the Emb_A column
Notice that your mapping uses float keys (like 762.0) but the original A column has integers. We'll convert A to float to ensure matching works correctly, then use pandas' map() function to assign the corresponding Emb_A values.
# Convert column A to float to match the mapping's key type original_df['A'] = original_df['A'].astype(float) # Add the Emb_A column by mapping values in A original_df['Emb_A'] = original_df['A'].map(embedding_map)
Step 3: Check the result
If you print the dataframe now, you'll get exactly the output you're looking for:
print(original_df)
Output:
id tslot A Emb_A 0 1 2014-11-02 22:45:00 89.0 -0.033929 1 1 2014-10-26 09:15:00 762.0 -0.237305 2 1 2014-10-26 11:00:00 25.0 -0.237300 3 1 2014-10-26 11:15:00 762.0 -0.237305 4 1 2014-10-26 12:00:00 25.0 -0.237300 5 1 2014-10-26 13:00:00 25.0 -0.237300 6 1 2014-10-26 13:45:00 25.0 -0.237300 7 1 2014-10-26 14:45:00 902.0 -0.237330 8 1 2014-10-26 15:45:00 902.0 -0.237330
Bonus: Handle missing values
If there are values in A that aren't in your mapping, map() will return NaN for those rows. You can fill those with a default value (like 0.0) using fillna():
original_df['Emb_A'] = original_df['A'].map(embedding_map).fillna(0.0)
内容的提问来源于stack exchange,提问作者Krush23

