如何在Polars中使用read_csv处理行1后列数增加的CSV?
用Polars读取列数不规则的CSV文件
问题场景
CSV文件中行的列数不固定,部分行仅1列,其余行有2列。使用Polars默认的read_csv方法时,只会识别出1列,且无法通过skiprows参数跳过单列行(因为单列行的位置不固定)。需要实现类似Pandas中names参数的效果,强制读取指定列数,不足列自动填充null。
解决方案
使用Polars的read_csv方法时,通过columns参数明确指定要读取的列名(列名数量对应目标列数),Polars会自动将列数不足的行的缺失列填充为null。
代码示例
import polars as pl data = b""" Data Date,A Time,B """.strip() # 指定2列的列名,强制读取2列 df = pl.read_csv(data, has_header=False, columns=["column_1", "column_2"]) print(df)
输出结果
shape: (3, 2) ┌──────────┬──────────┐ │ column_1 ┆ column_2 │ │ --- ┆ --- │ │ str ┆ str │ ╞══════════╪══════════╡ │ Data ┆ null │ │ Date ┆ A │ │ Time ┆ B │ └──────────┴──────────┘
动态适配最大列数(可选)
如果无法提前确定目标列数,可以先预扫描文件获取最大列数,再生成对应数量的列名:
import polars as pl from io import BytesIO def get_max_columns(data): max_cols = 0 with BytesIO(data) as f: for line in f: line = line.decode().strip() if not line: continue cols = len(line.split(",")) if cols > max_cols: max_cols = cols return max_cols data = b""" Data Date,A Time,B,C """.strip() max_cols = get_max_columns(data) columns = [f"column_{i+1}" for i in range(max_cols)] df = pl.read_csv(data, has_header=False, columns=columns) print(df)
内容的提问来源于stack exchange,提问作者Josh Y.
相关产品推荐
相关产品推荐

