Pandas CSV数据清理:修复列值错位问题
Fixing Column Misalignment from Missing Values in Pandas
Hey, I've got you covered on fixing this column misalignment issue! Here's how you can implement the data cleaning logic you need, building on the CSV loading code you already have:
Step-by-Step Solution
First, here's the complete code that will handle the correction:
import pandas as pd # Your existing data loading code df = pd.read_csv('data.csv', dtype={ 'date': str, 'tap': str, 'time': str, 'count': float }) # Create a mask to identify rows with invalid 'tap' values invalid_tap_mask = ~df['tap'].isin(['on', 'off']) # Correct the column values for these misaligned rows df.loc[invalid_tap_mask, 'count'] = df.loc[invalid_tap_mask, 'time'] df.loc[invalid_tap_mask, 'time'] = df.loc[invalid_tap_mask, 'tap'] df.loc[invalid_tap_mask, 'tap'] = 'N/A'
How This Works:
- Mask Creation: The
invalid_tap_maskvariable is a boolean series whereTruemarks rows where the 'tap' value is neither 'on' nor 'off' (the~operator reverses the result ofisin()). - Column Correction: Using
df.loc[mask, column], we target only the problematic rows and shift values appropriately:- Move the original 'time' value into the 'count' column
- Move the original 'tap' value into the 'time' column
- Set the 'tap' column to 'N/A' to indicate the missing value
Verify the Fix
To double-check that everything worked as expected, you can print out the modified rows with:
print(df[invalid_tap_mask])
内容的提问来源于stack exchange,提问作者Ken
相关产品推荐
相关产品推荐

