You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何借助.names文件为Pandas导入的UCI auto-mpg数据集设置列名?

解决方案:为Auto-MPG数据集设置列名

步骤1:提取列名

auto-mpg.names是数据集的说明文档,其中明确列出了9个字段的名称,整理成列表如下:

column_names = [
    "mpg",
    "cylinders",
    "displacement",
    "horsepower",
    "weight",
    "acceleration",
    "model_year",
    "origin",
    "car_name"
]

步骤2:修改读取代码

在pd.read_csv中添加names参数传入列名列表,同时设置header=None(因为原始数据没有自带表头,避免把第一行数据误判为表头):

import pandas as pd

column_names = [
    "mpg",
    "cylinders",
    "displacement",
    "horsepower",
    "weight",
    "acceleration",
    "model_year",
    "origin",
    "car_name"
]

data_auto = pd.read_csv(
    "https://archive.ics.uci.edu/ml/machine-learning-databases/auto-mpg/auto-mpg.data-original",
    comment="#",
    sep=r"\s+",
    names=column_names,
    header=None
)

# 验证结果
print(data_auto.head())

关键说明

  • sep=r"\s+"的作用:匹配任意数量的空白字符(空格、制表符),适配数据集中字段间的不规则分隔。
  • header=None必须设置:原始数据第一行是实际数据,不是表头,该参数告诉Pandas不要将第一行当作列名处理。

内容的提问来源于stack exchange,提问作者Zofia Smoleń

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 08:35:21