使用R读取GDC下载的tar.gz文件后无法查看表格,求指导
Hey there! Let's get that tar.gz file's content into a usable table in R. The issue right now is that you've only created a connection to the gzipped file, but this is actually a tar archive (a collection of files wrapped together and then gzipped). So we need to handle the tar layer first to access the actual table files inside.
Here are two straightforward methods to do this:
Method 1: Read files directly from the tar.gz (no need to extract to disk)
First, let's check what files are inside the archive:
# Open the gzipped connection gz_conn <- gzfile("D:/New folder/gdc_download_20191030_052506.304900.tar.gz", open = "rb") # List all files in the tar archive archive_files <- untar(gz_conn, list = TRUE) # Close the connection when done close(gz_conn) # Print the list of files to see what's inside print(archive_files)
Once you know which file you want (e.g., something like sample_data.tsv), you can extract and read it directly:
# Re-open the connection gz_conn <- gzfile("D:/New folder/gdc_download_20191030_052506.304900.tar.gz", open = "rb") # Extract the target file to a temporary directory untar(gz_conn, files = "sample_data.tsv", exdir = tempdir()) close(gz_conn) # Read the table (use read.delim for TSV files, common in GDC data) data_table <- read.delim(file.path(tempdir(), "sample_data.tsv"), header = TRUE) # Check the first few rows head(data_table)
Method 2: Extract the entire archive first, then read files
If you prefer to have the files saved locally for easy access:
# Extract all contents to a new folder untar("D:/New folder/gdc_download_20191030_052506.304900.tar.gz", exdir = "D:/New folder/gdc_extracted") # List the extracted files to find your table list.files("D:/New folder/gdc_extracted") # Read the table (adjust the file name and read function to match your file type) data_table <- read.delim("D:/New folder/gdc_extracted/your_target_file.tsv", header = TRUE) head(data_table)
Bonus: Batch-read multiple files (common with GDC data)
GDC downloads often include multiple sample files. If you need to combine them into one table:
# Get paths to all TSV files in the extracted folder file_paths <- list.files("D:/New folder/gdc_extracted", pattern = "\\.tsv$", full.names = TRUE) # Read all files into a list of data frames all_samples <- lapply(file_paths, read.delim, header = TRUE) # Combine into a single data frame (only if all files have the same structure) combined_data <- do.call(rbind, all_samples) head(combined_data)
A quick note: GDC data is typically tab-separated (TSV), so read.delim (which defaults to tab separators) works better than read.csv here. If your file uses a different separator, just adjust the sep argument in the read function.
内容的提问来源于stack exchange,提问作者Dev P

