在R中读取多份TSV文件并合并为带行列名的矩阵/数据框问题
It sounds like the main issue you’re hitting is that tibbles (the default data frame type in the tidyverse) don’t natively support row names—so when combining your data, those row names get dropped unless you explicitly preserve them as a column first. Let’s fix that with an adjusted workflow using purrr, readr, and dplyr:
Step 1: Prepare Full File Paths
First, let’s complete your file path generation (I’ll assume each sample has a corresponding TSV named like {sample}.tsv):
library(tidyverse) TARGET_FOLDER <- "your_target_folder_here" # Replace with your actual folder path samples <- c("a","b","c","d") # Generate full paths to each TSV file file_paths <- samples %>% map_chr(~ file.path(getwd(), TARGET_FOLDER, paste0(.x, ".tsv")))
Step 2: Read Files with Row Names Preserved
When reading each TSV, we’ll convert row names into a dedicated column (e.g., rowname) since tibbles don’t store row names natively. I’ll assume your TSVs have row names in the first column (adjust row_names if your setup differs):
# Custom function to read TSV and keep row names as a column read_tsv_with_rownames <- function(file_path, sample_name) { read_tsv(file_path, row_names = TRUE) %>% rownames_to_column(var = "rowname") %>% # Convert row names to a proper column select(rowname, value = everything()) %>% # Extract the data column (temp name) rename(!!sample_name := value) # Rename column to match the sample ID } # Apply the function to all files sample_data <- file_paths %>% map2(samples, read_tsv_with_rownames)
Step 3: Merge All Columns
Now we can merge all the data frames using the shared rowname column to keep everything aligned:
# Combine all sample data into one data frame combined_df <- sample_data %>% reduce(full_join, by = "rowname") %>% column_to_rownames(var = "rowname") # Convert the rowname column back to row names if needed
Step 4: Convert to Matrix (Optional)
If you need a matrix instead of a data frame:
combined_matrix <- as.matrix(combined_df)
Key Notes
- By explicitly converting row names to a column with
rownames_to_column, we ensure they’re carried through every tidyverse operation. - Using
map2lets us pass both the file path and sample name to our reading function, so we can rename columns to match samples immediately. reduce(full_join)combines all data frames using therownamecolumn, preserving all rows even if some samples are missing data (useinner_joininstead if you only want rows present in all samples).
If your TSVs have a specific column you need to extract (not the first data column), adjust the select step to target that column by name (e.g., select(rowname, value = your_target_column)).
内容的提问来源于stack exchange,提问作者ftabaro

