Pandas DataFrame列长度不一致时如何将列值合并至单行
Hey there! Let's work through this problem where your ID, W_Weight, and Class columns don't line up in length, plus you noted that each ID should correspond to exactly one Class value (like ID 0 maps to Class 1.0, ID 4 maps to Class 5.0). First, I'll clean up the raw data you shared into a clear table since the original input had some formatting mess:
| ID | W_Weight | Class |
|---|---|---|
| 0 | 0.255265 | 1.0 |
| 0 | 0.273844 | 1.0 |
| 0 | 0.351219 | 1.0 |
| 0 | 0.262033 | 1.0 |
| 0 | 0.351219 | 5.0 |
| 0 | 0.258109 | 1.0 |
| 0 | 0.296328 | 5.0 |
| 0 | 0.351219 | 1.0 |
| 0 | 0.301208 | 1.0 |
| 0 | 0.273844 | 1.0 |
| 0 | 0.317767 | 1.0 |
| 1 | 0.299451 | 1.0 |
| 1 | 0.327183 | 5.0 |
| 1 | 0.391577 | 1.0 |
| 1 | 0.272526 | 1.0 |
| 1 | 0.41 | (Missing) |
Step 1: Align Column Lengths & Fill Gaps
First, we need to make sure all three columns have the same number of rows. If the Class column is missing values (like the last row in your data), we'll mark those gaps with NaN temporarily, then fix them using the ID-Class mapping rule.
Here's how to do this with Pandas (the go-to tool for this kind of data cleanup):
import pandas as pd # Convert your raw data into a DataFrame data = [ [0, 0.255265, 1.0], [0, 0.273844, 1.0], [0, 0.351219, 1.0], [0, 0.262033, 1.0], [0, 0.351219, 5.0], [0, 0.258109, 1.0], [0, 0.296328, 5.0], [0, 0.351219, 1.0], [0, 0.301208, 1.0], [0, 0.273844, 1.0], [0, 0.317767, 1.0], [1, 0.299451, 1.0], [1, 0.327183, 5.0], [1, 0.391577, 1.0], [1, 0.272526, 1.0], [1, 0.41, None] ] df = pd.DataFrame(data, columns=["ID", "W_Weight", "Class"])
Step 2: Fix Class Values Using ID Mapping
Since each ID should link to exactly one Class, we have two solid ways to fix mismatched or missing Class values:
Option 1: Manual Mapping (If You Know Exact ID-Class Pairs)
If you already have a definitive list of which ID maps to which Class, create a dictionary and apply it to the DataFrame:
# Define your official ID-to-Class mapping id_class_map = { 0: 1.0, 1: 1.0, # Adjust this to match your actual correct value for ID 1 4: 5.0 # Add more ID-Class pairs as needed } # Replace all Class values with the correct one for their ID df["Class"] = df["ID"].map(id_class_map)
Option 2: Auto-Fix with Mode (If Most Rows Have the Correct Class)
If only a few rows have wrong Class values (like ID 0 mostly has 1.0 but a couple rows show 5.0), we can use the most frequent Class value for each ID:
# Calculate the most common (mode) Class for each ID id_mode_class = df.groupby("ID")["Class"].agg(pd.Series.mode).to_dict() # Handle cases where multiple modes exist (pick the first one) id_mode_class = {k: v[0] if isinstance(v, list) else v for k, v in id_mode_class.items()} # Update the Class column df["Class"] = df["ID"].map(id_mode_class)
Step 3: Verify the Fix
Finally, double-check that everything is working as expected:
# Confirm all columns have the same length print("Columns are same length?", len(df["ID"]) == len(df["W_Weight"]) == len(df["Class"])) # Confirm each ID has only one Class value print("Each ID has unique Class?", df.groupby("ID")["Class"].nunique().max() == 1)
This will resolve the column length mismatch and ensure every row's Class value matches the correct one for its ID.
内容的提问来源于stack exchange,提问作者Amjad Dahlawi

