You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用csvwriter写入CSV文件致程序周末崩溃,求解决方案

持续写入CSV文件停滞问题的排查与解决

问题背景

我在一台电脑上运行两个Python程序,共打开5个CSV文件,通过writer.writerow()每秒向每个文件写入约1行数据,程序计划持续运行数周。周末回来后发现两个程序无报错崩溃,且部分CSV文件在崩溃前很久就已经停止写入,怀疑是内存溢出导致的问题。以下是我用于操作CSV文件的类:

class csv_file:
    
    ROW_HEADER = ["Column 1", "Column 2", "Column 3"]
    
    def __init__(self, name: str, path: str) -> None:
        self.file_name = f"{name}.csv"
        self.file_path = path
        self.file = None
        self.lock = threading.RLock()
        self.start_time = time.time()

    def create_file(self) -> None:
        '''
        Creates a csv file with the name and path specified in the constructor.
        Appends the row header to the file.
        '''
        with self.lock:
            os.makedirs(self.file_path, exist_ok=True) # Create the directory if it doesn't exist (nothing happens if it does)
            self.file = open(os.path.join(self.file_path, self.file_name), 'w')
            self.writer = csv.writer(self.file)
            self.writer.writerow(self.ROW_HEADER)

    def _create_row(self, row:list) -> None:
        '''
        Adds a row to the csv file. The row must be the same
        size as the row header.
        '''
        with self.lock:
            self.writer.writerow(row)

    def add_row(self, val1:float, val2:float, val3:float) -> None:
        with self.lock:
            row = []
            row.append(val1)
            row.append(val2)
            row.append(val3)
            self._create_row(row)

排查与解决建议

1. 核心原因:文件缓冲区积压

Python文件对象默认使用缓冲机制,调用writer.writerow()后数据不会立刻写入磁盘,而是暂存在内存缓冲区。长时间运行后,缓冲区持续积压,不仅会占用大量内存,还可能因为缓冲区未触发自动刷新,导致写入停滞,甚至程序无响应。

2. 优先修复:强制刷新缓冲区

每次写入后手动刷新缓冲区,确保数据立刻落盘,这是最直接的解决方案:
修改_create_row方法:

def _create_row(self, row:list) -> None:
    with self.lock:
        self.writer.writerow(row)
        # 刷新缓冲区到内存
        self.file.flush()
        # 可选:强制操作系统将内存数据写入磁盘,适合对数据完整性要求极高的场景
        # import os
        # os.fsync(self.file.fileno())

3. 备选方案:避免长期持有文件句柄

如果程序允许,可以每次写入时打开文件,写入后立即关闭,彻底避免缓冲区积压问题。但这种方式在高频率写入时会有性能损耗,需要根据实际场景权衡:

def _create_row(self, row:list) -> None:
    with self.lock:
        full_path = os.path.join(self.file_path, self.file_name)
        # 用追加模式打开,newline=''避免csv写入时出现空行
        with open(full_path, 'a', newline='') as f:
            writer = csv.writer(f)
            writer.writerow(row)

注意:需要确保create_file方法已经写入表头,后续用追加模式写入数据。

4. 内存监控与排查

  • 添加内存监控:用psutil模块定期记录内存占用,确认是否真的是内存溢出导致的问题:
    import psutil
    import time
    
    def log_memory_usage():
        process = psutil.Process()
        mem_rss = process.memory_info().rss / (1024 * 1024)
        print(f"[{time.strftime('%Y-%m-%d %H:%M:%S')}] 内存占用:{mem_rss:.2f} MB")
        # 也可以把日志写入文件留存
    
    可以把这个函数放在写入循环中,每隔一段时间调用一次。
  • 检查其他内存泄漏点:排查程序中是否有其他对象(比如全局列表、字典)在不断累积数据,没有及时清理,这也是长期运行程序内存溢出的常见原因。

5. 稳定性增强

  • 添加异常捕获:在写入逻辑中加入异常处理,避免程序静默崩溃,同时记录错误信息:
    def _create_row(self, row:list) -> None:
        with self.lock:
            try:
                self.writer.writerow(row)
                self.file.flush()
            except IOError as e:
                # 可以把错误写入日志文件,方便后续排查
                print(f"[{time.strftime('%Y-%m-%d %H:%M:%S')}] CSV写入错误:{str(e)}")
    
  • 定期重启文件句柄:如果必须长期持有文件句柄,可以每隔一段时间(比如每天)重新打开一次文件,避免系统层面的文件句柄异常导致写入停滞。

内容的提问来源于stack exchange,提问作者jlaufer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 19:54:50