如何在R中加载.data格式数据集?如何确定正确编码?
正确加载UCI蘑菇数据集到R的方法
你提到的.data后缀文件本质是逗号分隔的纯文本文件(ASCII编码),之前的代码存在几个问题:
- 错误指定
UTF-16编码,该文件为标准ASCII格式,无需特殊编码设置 - 设置
header=T,但数据集本身没有表头行 dec=","参数多余,数据中全是分类变量,没有小数类型字段
下面是两种靠谱的加载方式:
方法1:用read.csv(最简便)
read.csv默认处理逗号分隔的文件,完美适配这个数据集:
# 直接读取远程数据集 mushroom_data <- read.csv("https://archive.ics.uci.edu/ml/machine-learning-databases/mushroom/agaricus-lepiota.data", header = FALSE) # 可选:给列添加对应表头(参考UCI数据集的属性定义) col_names <- c("class", "cap-shape", "cap-surface", "cap-color", "bruises", "odor", "gill-attachment", "gill-spacing", "gill-size", "gill-color", "stalk-shape", "stalk-root", "stalk-surface-above-ring", "stalk-surface-below-ring", "stalk-color-above-ring", "stalk-color-below-ring", "veil-type", "veil-color", "ring-number", "ring-type", "spore-print-color", "population", "habitat") colnames(mushroom_data) <- col_names
方法2:用read.table手动配置参数
如果习惯用read.table,修正参数即可:
mushroom_data <- read.table("https://archive.ics.uci.edu/ml/machine-learning-databases/mushroom/agaricus-lepiota.data", sep = ",", header = FALSE, stringsAsFactors = FALSE) # 可选:避免自动转为因子类型,按需设置
加载完成后可以用head(mushroom_data)查看前几行,确认数据是否正确读取。
内容的提问来源于stack exchange,提问作者Dieu94
相关产品推荐
相关产品推荐

