csv.writerow写入行丢失问题排查:老旧机械硬盘设备下大行数CSV写入异常
Your core code logic isn't inherently wrong—it works perfectly on fast storage like SSDs/NVMe because those devices handle write buffering and IO operations much more reliably. The row loss on HDDs comes down to how Python interacts with slower storage, unhandled edge cases, and buffering behavior. Here's a breakdown of the key issues and solutions:
Key Causes of Row Loss on HDDs
Unflushed File Buffers
Python uses in-memory buffering for file operations to optimize performance. On fast SSDs, the buffer gets flushed to disk quickly when the file is closed (via thewithstatement). But on slower HDDs, especially with limited RAM (like your 4GB test machine), the buffer might not fully write to disk before the program exits or moves to the next file—even though thewithblock should handle this, HDDs can have slower disk cache flushes that lead to data being stuck in transit.Conflicting Line Terminator Settings
You've setlineterminator='\n'incsv.writerandnewline=''inopen(). While this works on some systems, on Windows (where your HDD tests are running), this can cause inconsistent line ending handling. The CSV module expects to manage line endings whennewline=''is set, and overridinglineterminatormight lead to some rows not being recognized as valid lines when the file is read back (or even not written properly in the first place).Silent IO Errors
Older HDDs are more prone to minor IO errors (e.g., slow seek times, temporary read/write glitches) that might not crash your program but can cause individual rows to fail to write. Your current code doesn't catch these errors, so they happen silently and result in missing rows.
Fixes to Try
1. Force Buffer Flushes
Manually flush the buffer during writes to ensure data gets pushed to the disk immediately, instead of waiting for the buffer to fill up or the file to close. You can either use line buffering or flush after batches of rows:
for col in data: # Use line buffering with buffering=1, or flush after each row/batch with open('output/{}.csv'.format(col), mode, encoding='utf-8', newline='', buffering=1) as f: writer = csv.writer(f, lineterminator='\n') if mode == 'w': writer.writerow(headers) for row in data[col]: writer.writerow(row) f.flush() # Force data to disk right away
For better performance, you can flush every N rows (e.g., every 1000 rows) instead of every single row.
2. Let the CSV Module Handle Line Endings
On Windows, remove the lineterminator='\n' parameter from csv.writer. When you set newline='' in open(), the CSV module automatically uses the correct line endings for your system (\r\n on Windows), which avoids line ending mismatches that can cause row counting issues:
with open('output/{}.csv'.format(col), mode, encoding='utf-8', newline='') as f: writer = csv.writer(f) # No lineterminator override if mode == 'w': writer.writerow(headers) for row in data[col]: writer.writerow(row)
3. Add Error Handling and Logging
Catch IO errors during writing to identify if specific rows or files are failing, and log the issues for debugging:
import logging logging.basicConfig(filename='csv_write_errors.log', level=logging.INFO) for col in data: try: with open('output/{}.csv'.format(col), mode, encoding='utf-8', newline='') as f: writer = csv.writer(f) row_count = 0 if mode == 'w': writer.writerow(headers) row_count += 1 for row in data[col]: try: writer.writerow(row) row_count += 1 except IOError as e: logging.error(f"Failed to write row {row_count} for {col}.csv: {str(e)}") logging.info(f"Completed writing {row_count} rows for {col}.csv") except Exception as e: logging.error(f"Failed to process {col}.csv entirely: {str(e)}")
This will help you confirm if the row loss is due to specific IO errors or just buffering.
4. Batch Writes with Verification
Write rows in batches and verify the number of rows written matches the expected count. For example:
batch_size = 1000 for col in data: expected_rows = len(data[col]) + (1 if mode == 'w' else 0) with open('output/{}.csv'.format(col), mode, encoding='utf-8', newline='') as f: writer = csv.writer(f) if mode == 'w': writer.writerow(headers) written = 0 batch = [] for row in data[col]: batch.append(row) if len(batch) >= batch_size: writer.writerows(batch) f.flush() written += len(batch) batch = [] if batch: writer.writerows(batch) f.flush() written += len(batch) # Verify the count if mode == 'w': written += 1 # Add header row print(f"Expected {expected_rows} rows for {col}.csv, wrote {written}")
This helps you catch discrepancies immediately.
Final Notes
Your original code works on fast storage because the write operations complete quickly and buffers are flushed reliably. On HDDs, the slower IO and potential for minor errors expose gaps in the default buffering and error handling. By adjusting these settings, you'll get consistent row counts across all storage types.
内容的提问来源于stack exchange,提问作者dinnertoast

