如何处理1500万行大型CSV文件:分块删最后一行或直接定位删除?
处理大型CSV文件删除最后一行的两种方案
一、无需读取整个文件直接删除最后一行
对于1500万行的超大CSV,直接操作文件末尾是最高效的方案,完全不需要加载整个文件到内存。
Python实现方案
利用文件指针从后往前扫描,定位到倒数第二行的结束位置后截断文件:
def remove_last_line(file_path): with open(file_path, 'r+b') as f: # 移动到文件末尾前的起始位置 f.seek(0, 2) pos = f.tell() - 1 # 从后往前寻找倒数第二个换行符 while pos > 0 and f.read(1) != b'\n': pos -= 1 f.seek(pos) # 截断文件到目标位置 if pos > 0: f.seek(pos) f.truncate() else: # 处理文件只有一行的情况 f.truncate(0)
注意:如果你的CSV行结尾是\r\n(Windows格式),需要把判断条件改为f.read(2) == b'\r\n',并调整指针移动的步长。
Shell命令快速实现(Linux/macOS)
用系统命令可以快速完成,适合不需要写代码的场景:
# 计算最后一行的字节长度,然后截断文件 truncate -s $(($(stat -c %s filename.csv) - $(tail -n1 filename.csv | wc -c))) filename.csv
注意:如果文件最后一行末尾有换行符,需要在计算时额外减去1个字节。
二、分块方式删除最后一行
如果因为环境限制无法直接操作文件指针,或者需要在处理过程中校验行内容,可以用分块读取的方式,将除最后一行外的内容写入临时文件,再替换原文件。
Python实现方案
import os def remove_last_line_chunked(file_path, chunk_size=1024*1024): # 1MB分块,可按需调整 temp_file = f"{file_path}.tmp" with open(file_path, 'r', encoding='utf-8') as f_in, open(temp_file, 'w', encoding='utf-8') as f_out: last_chunk = "" while True: chunk = f_in.read(chunk_size) if not chunk: break # 合并当前块与上一块的剩余内容,按行拆分 combined = last_chunk + chunk lines = combined.splitlines(True) # 保留每行的换行符 if len(lines) > 1: # 写入除最后一行外的所有内容 f_out.write(''.join(lines[:-1])) last_chunk = lines[-1] else: last_chunk = combined # 处理最后剩余的内容,跳过最后一行 if last_chunk: final_lines = last_chunk.splitlines() if len(final_lines) > 1: f_out.write('\n'.join(final_lines[:-1]) + '\n') # 替换原文件 os.replace(temp_file, file_path)
注意:这种方法需要额外的磁盘空间存储临时文件,分块大小可根据可用内存调整,越大处理效率越高。
内容的提问来源于stack exchange,提问作者Sajjad Khan
相关产品推荐
相关产品推荐

