You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

csv.writerow写入行丢失问题排查:老旧机械硬盘设备下大行数CSV写入异常

Why Your CSV Writing Loses Rows on HDDs (and How to Fix It)

Your core code logic isn't inherently wrong—it works perfectly on fast storage like SSDs/NVMe because those devices handle write buffering and IO operations much more reliably. The row loss on HDDs comes down to how Python interacts with slower storage, unhandled edge cases, and buffering behavior. Here's a breakdown of the key issues and solutions:

Key Causes of Row Loss on HDDs

  1. Unflushed File Buffers
    Python uses in-memory buffering for file operations to optimize performance. On fast SSDs, the buffer gets flushed to disk quickly when the file is closed (via the with statement). But on slower HDDs, especially with limited RAM (like your 4GB test machine), the buffer might not fully write to disk before the program exits or moves to the next file—even though the with block should handle this, HDDs can have slower disk cache flushes that lead to data being stuck in transit.

  2. Conflicting Line Terminator Settings
    You've set lineterminator='\n' in csv.writer and newline='' in open(). While this works on some systems, on Windows (where your HDD tests are running), this can cause inconsistent line ending handling. The CSV module expects to manage line endings when newline='' is set, and overriding lineterminator might lead to some rows not being recognized as valid lines when the file is read back (or even not written properly in the first place).

  3. Silent IO Errors
    Older HDDs are more prone to minor IO errors (e.g., slow seek times, temporary read/write glitches) that might not crash your program but can cause individual rows to fail to write. Your current code doesn't catch these errors, so they happen silently and result in missing rows.

Fixes to Try

1. Force Buffer Flushes

Manually flush the buffer during writes to ensure data gets pushed to the disk immediately, instead of waiting for the buffer to fill up or the file to close. You can either use line buffering or flush after batches of rows:

for col in data: 
    # Use line buffering with buffering=1, or flush after each row/batch
    with open('output/{}.csv'.format(col), mode, encoding='utf-8', newline='', buffering=1) as f: 
        writer = csv.writer(f, lineterminator='\n') 
        if mode == 'w': 
            writer.writerow(headers) 
        for row in data[col]: 
            writer.writerow(row)
            f.flush()  # Force data to disk right away

For better performance, you can flush every N rows (e.g., every 1000 rows) instead of every single row.

2. Let the CSV Module Handle Line Endings

On Windows, remove the lineterminator='\n' parameter from csv.writer. When you set newline='' in open(), the CSV module automatically uses the correct line endings for your system (\r\n on Windows), which avoids line ending mismatches that can cause row counting issues:

with open('output/{}.csv'.format(col), mode, encoding='utf-8', newline='') as f: 
    writer = csv.writer(f)  # No lineterminator override
    if mode == 'w': 
        writer.writerow(headers) 
    for row in data[col]: 
        writer.writerow(row)

3. Add Error Handling and Logging

Catch IO errors during writing to identify if specific rows or files are failing, and log the issues for debugging:

import logging
logging.basicConfig(filename='csv_write_errors.log', level=logging.INFO)

for col in data: 
    try:
        with open('output/{}.csv'.format(col), mode, encoding='utf-8', newline='') as f: 
            writer = csv.writer(f)
            row_count = 0
            if mode == 'w': 
                writer.writerow(headers)
                row_count += 1
            for row in data[col]: 
                try:
                    writer.writerow(row)
                    row_count += 1
                except IOError as e:
                    logging.error(f"Failed to write row {row_count} for {col}.csv: {str(e)}")
            logging.info(f"Completed writing {row_count} rows for {col}.csv")
    except Exception as e:
        logging.error(f"Failed to process {col}.csv entirely: {str(e)}")

This will help you confirm if the row loss is due to specific IO errors or just buffering.

4. Batch Writes with Verification

Write rows in batches and verify the number of rows written matches the expected count. For example:

batch_size = 1000
for col in data: 
    expected_rows = len(data[col]) + (1 if mode == 'w' else 0)
    with open('output/{}.csv'.format(col), mode, encoding='utf-8', newline='') as f: 
        writer = csv.writer(f) 
        if mode == 'w': 
            writer.writerow(headers) 
        written = 0
        batch = []
        for row in data[col]: 
            batch.append(row)
            if len(batch) >= batch_size:
                writer.writerows(batch)
                f.flush()
                written += len(batch)
                batch = []
        if batch:
            writer.writerows(batch)
            f.flush()
            written += len(batch)
        # Verify the count
        if mode == 'w':
            written += 1  # Add header row
        print(f"Expected {expected_rows} rows for {col}.csv, wrote {written}")

This helps you catch discrepancies immediately.

Final Notes

Your original code works on fast storage because the write operations complete quickly and buffers are flushed reliably. On HDDs, the slower IO and potential for minor errors expose gaps in the default buffering and error handling. By adjusting these settings, you'll get consistent row counts across all storage types.

内容的提问来源于stack exchange,提问作者dinnertoast

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 10:19:08