如何用Pandas高效读取文本文件的首行、倒数第二及倒数第一行
解决大文件中读取指定行(含负索引)的高效方案
方法一:基于Pandas的索引转换方案
Pandas的skiprows不支持负索引,我们可以先通过低内存方式获取文件总行数,再将负索引转为正索引:
- 快速统计文件行数(不加载整个文件):
def get_line_count(file_path): with open(file_path, 'r', encoding='utf-8') as f: # 逐行计数,内存占用极低 return sum(1 for _ in f)
- 转换索引并读取目标行:
import pandas as pd txt_path = "你的文件路径.txt" rows_wanted = [0, -2, -1] # 获取总行数 total_lines = get_line_count(txt_path) # 转换负索引为正索引 target_rows = [] for idx in rows_wanted: if idx < 0: target_rows.append(total_lines + idx) else: target_rows.append(idx) # 仅读取目标行 data = pd.read_csv(txt_path, skiprows=lambda x: x not in target_rows)
方法二:直接文件操作(更高效,适合超大文件)
如果文件体积极大,可跳过全量遍历,仅读取第一行和最后两行:
import os import pandas as pd txt_path = "你的文件路径.txt" # 读取第0行 with open(txt_path, 'r', encoding='utf-8') as f: first_line = f.readline().strip() # 读取最后n行的工具函数 def get_last_n_lines(file_path, n): with open(file_path, 'rb') as f: # 从文件末尾反向查找换行符 f.seek(-2, os.SEEK_END) newline_count = 0 while newline_count < n: try: f.seek(-2, os.SEEK_CUR) except OSError: f.seek(0) break if f.read(1) == b'\n': newline_count += 1 last_lines = f.read().decode('utf-8').splitlines() return last_lines[-n:] if len(last_lines) >= n else last_lines # 获取最后两行 last_two_lines = get_last_n_lines(txt_path, 2) # 合并目标行并转为DataFrame all_target_lines = [first_line] + last_two_lines data = pd.DataFrame([line.split(',') for line in all_target_lines])
注意:若CSV文件含表头,需根据实际结构调整索引对应关系。
内容的提问来源于stack exchange,提问作者Mika R.
相关产品推荐
相关产品推荐

