在Python中截断大CSV文件:仅保留前n条记录
解决大CSV文件截断前n条记录的问题
Hey Leon, I totally feel your frustration with giant CSV files that won't fit into your laptop's memory—let's walk through several straightforward ways to save just the first n rows without loading the entire file.
方法1:用Pandas直接读取前n行(最简便)
Pandas的read_csv自带nrows参数,专门用来指定只读取前N行,完全不用加载整个10GB文件,内存友好得很:
import pandas as pd # 替换成你需要的目标行数 n = 10000 # 仅读取前n行数据 df_top_n = pd.read_csv("path/data.csv", nrows=n) # 保存到新文件,index=False避免生成额外的索引列 df_top_n.to_csv("path/top_n_data.csv", index=False)
这个方法最适合你已经习惯用Pandas的场景,代码简洁,操作高效。
方法2:逐块读取(适合超大n值)
如果n特别大,担心一次性读取n行还是会占用较多内存,可以用chunksize分块读取,直到收集够目标行数:
import pandas as pd n = 10000 chunk_size = 1000 # 每次读取1000行,可根据你的内存情况调整 collected_rows = [] total = 0 for chunk in pd.read_csv("path/data.csv", chunksize=chunk_size): if total + len(chunk) <= n: collected_rows.append(chunk) total += len(chunk) else: # 取剩余需要的行数并终止循环 collected_rows.append(chunk.iloc[:n - total]) total = n break # 合并所有块并保存 df_top_n = pd.concat(collected_rows, ignore_index=True) df_top_n.to_csv("path/top_n_data.csv", index=False)
方法3:Python内置文件操作(零依赖,极致省内存)
如果你不想依赖Pandas,直接用Python原生的文件读写功能,逐行处理,内存占用几乎可以忽略:
n = 10000 with open("path/data.csv", 'r') as in_file, open("path/top_n_data.csv", 'w') as out_file: # 先写入表头(如果你的CSV没有表头,删掉这两行即可) header = next(in_file) out_file.write(header) # 写入前n条数据行 for count, line in enumerate(in_file): if count < n: out_file.write(line) else: break
这个方法完全不需要任何第三方库,适合环境受限的场景。
方法4:命令行工具(最快最直接,适合Linux/macOS)
如果你用的是Linux或者macOS,直接用head命令就能搞定,连Python都不用写:
# 注意:如果CSV有表头,要取n条数据的话,需要取n+1行(包含表头) # 比如要10000条数据,就取10001行 head -n 10001 path/data.csv > path/top_n_data.csv # 如果没有表头,直接取n行即可 head -n 10000 path/data.csv > path/top_n_data.csv
这个方法速度最快,因为是系统级别的操作,完全绕开了Python的内存限制。
内容的提问来源于stack exchange,提问作者Leon
相关产品推荐
相关产品推荐

