You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用np.select处理后CSV输出文件行数与源文件不一致求助

Troubleshooting Missing Rows When Processing CSV with pandas and np.select

Hey Sam, let's break down why you're losing hundreds of rows when writing to file2—it's likely one of a few easy-to-miss issues with how you're reading, processing, or writing the data.

1. First: Verify if Rows Are Missing Before Processing

Before blaming np.select or the write step, check if the problem starts right when you read the file. Run these lines immediately after reading file1:

print(f"Original file1 row count: [手动统计file1的实际行数]")
print(f"Read DataFrame row count: {len(imu_gyr)}")

If the counts don't match here, the issue lies with pd.read_csv, not your processing:

  • Check index_col=0: If the first column of file1 isn't a unique, non-null index, using index_col=0 can cause pandas to drop rows or merge entries (especially if there are duplicate index values). Try removing this parameter and see if the row count matches:
    imu_gyr = pd.read_csv(file1)  # Omit index_col for now
    
  • Check for malformed lines: Add on_bad_lines='warn' to pd.read_csv to see if pandas is silently skipping invalid rows:
    imu_gyr = pd.read_csv(file1, index_col=0, on_bad_lines='warn')
    

2. Fix the np.select Syntax (And Simplify 16-bit Conversion)

Looking at your code, there's a syntax error in the np.select call—you have an extra closing bracket ] at the end. Even if Python didn't throw an error (maybe a typo when pasting), your 16-bit two's complement conversion can be simplified to avoid manual checks, which reduces room for bugs:

Instead of manually checking values >32767, let numpy handle the two's complement conversion directly:

# Combine LSB and MSB into a 16-bit unsigned integer
combined_values = imu_gyr["LSB"] + imu_gyr["MSB"] * 256
# Convert to signed 16-bit integer to auto-handle two's complement logic
signed_values = combined_values.astype(np.uint16).astype(np.int16)
# Calculate final Gyr_x value
imu_gyr["Gyr_x"] = signed_values / 1000

This does exactly the same logic as your np.select but is cleaner and less prone to mistakes.

3. Check CSV Writing Behavior

If the row count matches after processing but drops when writing, check these points:

  • Ensure no hidden filters: Double-check if you accidentally added a dropna() or row filter somewhere that you didn't mention. By default, to_csv doesn't drop rows with NaN values.
  • Verify output file integrity: Make sure File_Name_G() isn't returning a filename that's being overwritten by another process. You can temporarily add a timestamp to the filename to rule out accidental overwrites.

Full Corrected Code Example

Here's a cleaned-up version of your code incorporating the fixes above:

import pandas as pd
import numpy as np

file2 = File_Name_G()
# Read without index_col unless you're certain the first column is a unique index
imu_gyr = pd.read_csv(file1, on_bad_lines='warn')

# Simplified 16-bit two's complement conversion
combined_values = imu_gyr["LSB"] + imu_gyr["MSB"] * 256
imu_gyr["Gyr_x"] = combined_values.astype(np.uint16).astype(np.int16) / 1000

# Confirm row counts before and after writing
print(f"Rows before writing: {len(imu_gyr)}")
imu_gyr.to_csv(file2, index=False)  # index=False avoids writing a default index column
written_df = pd.read_csv(file2)
print(f"Rows after writing: {len(written_df)}")

By following these steps, you should be able to pinpoint exactly where the rows are going missing. Let me know if any of these resolve your issue!

内容的提问来源于stack exchange,提问作者Samboff

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 14:47:48