使用np.select处理后CSV输出文件行数与源文件不一致求助
Hey Sam, let's break down why you're losing hundreds of rows when writing to file2—it's likely one of a few easy-to-miss issues with how you're reading, processing, or writing the data.
1. First: Verify if Rows Are Missing Before Processing
Before blaming np.select or the write step, check if the problem starts right when you read the file. Run these lines immediately after reading file1:
print(f"Original file1 row count: [手动统计file1的实际行数]") print(f"Read DataFrame row count: {len(imu_gyr)}")
If the counts don't match here, the issue lies with pd.read_csv, not your processing:
- Check
index_col=0: If the first column offile1isn't a unique, non-null index, usingindex_col=0can cause pandas to drop rows or merge entries (especially if there are duplicate index values). Try removing this parameter and see if the row count matches:imu_gyr = pd.read_csv(file1) # Omit index_col for now - Check for malformed lines: Add
on_bad_lines='warn'topd.read_csvto see if pandas is silently skipping invalid rows:imu_gyr = pd.read_csv(file1, index_col=0, on_bad_lines='warn')
2. Fix the np.select Syntax (And Simplify 16-bit Conversion)
Looking at your code, there's a syntax error in the np.select call—you have an extra closing bracket ] at the end. Even if Python didn't throw an error (maybe a typo when pasting), your 16-bit two's complement conversion can be simplified to avoid manual checks, which reduces room for bugs:
Instead of manually checking values >32767, let numpy handle the two's complement conversion directly:
# Combine LSB and MSB into a 16-bit unsigned integer combined_values = imu_gyr["LSB"] + imu_gyr["MSB"] * 256 # Convert to signed 16-bit integer to auto-handle two's complement logic signed_values = combined_values.astype(np.uint16).astype(np.int16) # Calculate final Gyr_x value imu_gyr["Gyr_x"] = signed_values / 1000
This does exactly the same logic as your np.select but is cleaner and less prone to mistakes.
3. Check CSV Writing Behavior
If the row count matches after processing but drops when writing, check these points:
- Ensure no hidden filters: Double-check if you accidentally added a
dropna()or row filter somewhere that you didn't mention. By default,to_csvdoesn't drop rows with NaN values. - Verify output file integrity: Make sure
File_Name_G()isn't returning a filename that's being overwritten by another process. You can temporarily add a timestamp to the filename to rule out accidental overwrites.
Full Corrected Code Example
Here's a cleaned-up version of your code incorporating the fixes above:
import pandas as pd import numpy as np file2 = File_Name_G() # Read without index_col unless you're certain the first column is a unique index imu_gyr = pd.read_csv(file1, on_bad_lines='warn') # Simplified 16-bit two's complement conversion combined_values = imu_gyr["LSB"] + imu_gyr["MSB"] * 256 imu_gyr["Gyr_x"] = combined_values.astype(np.uint16).astype(np.int16) / 1000 # Confirm row counts before and after writing print(f"Rows before writing: {len(imu_gyr)}") imu_gyr.to_csv(file2, index=False) # index=False avoids writing a default index column written_df = pd.read_csv(file2) print(f"Rows after writing: {len(written_df)}")
By following these steps, you should be able to pinpoint exactly where the rows are going missing. Let me know if any of these resolve your issue!
内容的提问来源于stack exchange,提问作者Samboff

