使用Pandas读取CSV文件时保留原始行号及表头的方法
解决方法
要实现跳过指定行同时保留原始文件行号(1-based)和表头,有几种实用的方式:
方法1:先读取再调整索引(适合小文件)
如果文件不大,可以先完整读取,再重置索引为原始文件行号,最后筛选需要的行:
import pandas as pd # 完整读取CSV,默认表头为第1行(文件行号1) df = pd.read_csv("your_file.csv") # 将DataFrame索引替换为原始文件的行号(数据行对应文件行号从2开始) df.index = df.index + 2 # 筛选从文件第100行开始的内容 df = df.loc[100:]
方法2:用skiprows读取后手动设置索引(高效)
先通过skiprows跳过不需要的行,再手动将索引设置为原始文件行号:
import pandas as pd # 要开始读取的文件行号(1-based,表头是行1) start_file_row = 100 # 计算要跳过的数据行:文件行2到行99,对应pd.read_csv的0-based行号1到98 skip_rows = range(1, start_file_row - 1) # 读取时跳过指定行,保留表头 df = pd.read_csv("your_file.csv", skiprows=skip_rows) # 将索引设置为原始文件的行号序列 df.index = range(start_file_row, start_file_row + len(df))
方法3:分块读取(适合大文件)
如果文件过大,用分块读取避免内存占用过高:
import pandas as pd start_file_row = 100 chunksize = 1000 # 根据内存情况调整块大小 result_chunks = [] for chunk in pd.read_csv("your_file.csv", chunksize=chunksize): # 计算当前块对应的文件行号范围 chunk_file_start = chunk.index[0] + 2 chunk_file_end = chunk_file_start + len(chunk) - 1 # 只保留块中大于等于目标起始行的部分 if chunk_file_end >= start_file_row: mask = (chunk.index + 2) >= start_file_row result_chunks.append(chunk[mask]) # 合并所有块 df = pd.concat(result_chunks) # 设置索引为原始文件行号 df.index = range(start_file_row, start_file_row + len(df))
注意点
- 这里的文件行号是1-based:表头对应文件第1行,第一行数据对应文件第2行;如果你的文件表头不在第1行,需要调整
header参数和索引计算逻辑。 - 如果CSV本身自带行号列,直接用
index_col指定该行作为索引即可,无需手动计算。
内容的提问来源于stack exchange,提问作者M.Nemes
相关产品推荐
相关产品推荐

