R新手技术问询:如何读取三个CSV格式的数据集文件?
Hey there! Welcome to R—super stoked you're diving in. Reading those CSV files is totally straightforward, and I’ll break it down step by step so you don’t get stuck.
First, get your files sorted
First up, make sure you’ve downloaded train.csv, test.csv, and sampleSubmission.csv from your cloud storage and saved them all in the same folder on your computer. To make life easier, set R’s working directory to that folder—this way you don’t have to type out the full file path every time.
If you’re using RStudio, the easiest way is to go to Session > Set Working Directory > Choose Directory and pick your folder. If you prefer typing code, use this (replace the path with your actual folder location):
# Example paths: Windows = "C:/MyRProjects/KaggleData", Mac/Linux = "/home/yourname/RData" setwd("path/to/your/folder/with/csvs")
Option 1: Use base R (no extra packages needed)
R has a built-in function called read.csv() that works perfectly for most standard CSVs. Just run these lines to load each file:
# Load training set into a data frame train_data <- read.csv("train.csv") # Load test set test_data <- read.csv("test.csv") # Load sample submission file sample_sub <- read.csv("sampleSubmission.csv")
To check if it worked, type head(train_data) into the console—you’ll see the first 6 rows of your training data. If you just type train_data, it’ll print the whole dataset (or a preview if it’s big).
Option 2: Use readr for faster, smarter loading
If you’re working with larger datasets, or want more control over how R parses your columns, the readr package (part of the tidyverse toolkit) is a game-changer. First, install it (you only need to do this once):
install.packages("readr")
Then load it every time you start R:
library(readr)
Now use read_csv() (note the underscore instead of a dot) to load your files:
train_data <- read_csv("train.csv") test_data <- read_csv("test.csv") sample_sub <- read_csv("sampleSubmission.csv")
A cool perk here is that read_csv() tells you exactly what data types it assigned to each column—super helpful for catching mistakes early.
Quick fixes for common headaches
- Case sensitivity matters! R will throw an error if you type
Train.csvinstead oftrain.csv—double-check your file names. - If your CSV uses semicolons instead of commas (common in some regions), use
read.csv2()(base R) orread_csv2()(readr) instead. - If you get a "file not found" error, either double-check your working directory or use the full file path (like
read.csv("C:/MyRProjects/KaggleData/train.csv")).
You’re all set now—time to start exploring your data! If you hit any weird issues, feel free to ask for more help.
内容的提问来源于stack exchange,提问作者Akmal Masud

