You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R语言中读取多份CSV文件并添加文件名作为标识列的实现方案

Solution to Batch Read CSV Files and Add Source Filename Column in R

Hey there, let's walk through how to solve this problem efficiently using R's purrr package (for the map function) and dplyr for data manipulation. Here's a step-by-step breakdown:

Step 1: Load Required Packages

First, make sure you have these packages installed (run install.packages(c("purrr", "dplyr", "readr")) if you haven't already):

library(purrr)
library(dplyr)
library(readr) # For fast, tibble-friendly CSV reading

Step 2: Define Your File List

List out the CSV files you need to process:

file_paths <- c("file_1_october.csv", "file_2_november.csv")

Step 3: Use map_dfr to Read and Transform Files

We'll use map_dfr (a map variant that automatically binds results into a single data frame) to loop through each file, read it, and add the month column with the source filename:

final_dataset <- map_dfr(file_paths, function(current_file) {
  # Read the CSV file into a data frame
  raw_data <- read_csv(current_file)
  
  # Add the new month column, populated with the current file's name
  raw_data %>% 
    mutate(month = current_file)
})

What This Does

  • map_dfr iterates over every file in file_paths, running the anonymous function for each entry.
  • Inside the function, we first read the CSV content into a data frame.
  • mutate() creates a new month column where every row gets the exact filename it originated from.
  • Finally, map_dfr combines all individual data frames into one unified dataset with all rows and the new month column.

Base R Alternative (No External Packages)

If you prefer sticking to base R without tidyverse packages, here's the equivalent code:

file_paths <- c("file_1_october.csv", "file_2_november.csv")

final_dataset <- do.call(rbind, lapply(file_paths, function(current_file) {
  raw_data <- read.csv(current_file)
  raw_data$month <- current_file
  raw_data
}))

Check the Result

If you print final_dataset, you'll get exactly the structure you want:

# A tibble: 4 × 4
  name   age gender month              
  <chr> <dbl> <chr>  <chr>              
1 james    24 male   file_1_october.csv
2 Sue      21 female file_1_october.csv
3 Grey     24 male   file_2_november.csv
4 Juliet   21 female file_2_november.csv

内容的提问来源于stack exchange,提问作者John Karuitha

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 10:54:06