在R中处理WISE土壤数据库的结构化文本数据
Processing WISE Soil Database Text File into Two DataFrames
Hey there! Let's work through this WISE soil data file together—those messy, block-structured text files can be tricky, but we can break it down step by step. Here's a hands-on R solution to get the two dataframes you need:
Step 1: Read and Split the Data into Groups
First, we'll read the file and split it into individual data groups using the empty line separators:
# Read the text file dat <- readLines("test.txt") # Identify empty lines to split data groups empty_lines <- which(dat == "") # Create start/end indices for each group group_indices <- data.frame( start = c(1, empty_lines[-length(empty_lines)] + 1), end = c(empty_lines - 1, length(dat)) )
Step 2: Initialize Empty DataFrames
We'll set up empty dataframes to store our final results:
# DataFrame for site + SCOM information site_scom_df <- data.frame() # DataFrame for soil layer information (linked to LAT/LONG) soil_layer_df <- data.frame()
Step 3: Process Each Data Group
Now we'll loop through each group to extract and structure the data:
for (i in 1:nrow(group_indices)) { # Extract all rows for the current group current_group <- dat[group_indices$start[i]:group_indices$end[i]] ### Extract Site & SCOM Data (Rows 2-5) # Split header and values for site info (rows 2 & 3) site_headers <- strsplit(current_group[2], "\\s+")[[1]] site_values <- strsplit(current_group[3], "\\s+")[[1]] # Split header and values for SCOM info (rows 4 & 5) scom_headers <- strsplit(current_group[4], "\\s+")[[1]] scom_values <- strsplit(current_group[5], "\\s+")[[1]] # Combine into a single row dataframe site_scom_row <- as.data.frame(rbind(c(site_values, scom_values))) colnames(site_scom_row) <- c(site_headers, scom_headers) # Add to the site/SCOM dataframe site_scom_df <- rbind(site_scom_df, site_scom_row) ### Extract Soil Layer Data (Rows 6 onwards) # Combine split soil headers (row 6 + row 10 as per your description) soil_headers_part1 <- strsplit(current_group[6], "\\s+")[[1]] soil_headers_part2 <- strsplit(current_group[10], "\\s+")[[1]] soil_headers <- c(soil_headers_part1, soil_headers_part2) # Identify rows containing soil layer data (skip header rows 6 and 10) soil_data_rows <- c(7:9, 11:length(current_group)) soil_data_list <- lapply(current_group[soil_data_rows], function(row) strsplit(row, "\\s+")[[1]]) # Convert list to dataframe soil_layer_rows <- as.data.frame(do.call(rbind, soil_data_list)) colnames(soil_layer_rows) <- soil_headers # Link to the site's LAT/LONG for reference soil_layer_rows$LAT <- site_scom_row$LAT soil_layer_rows$LONG <- site_scom_row$LONG # Add to the soil layer dataframe soil_layer_df <- rbind(soil_layer_df, soil_layer_rows) }
Step 4: Clean Up Data Types
Finally, convert numeric columns to the correct type (adjust based on your actual fields):
# Convert numeric fields in site/SCOM dataframe site_scom_df$LAT <- as.numeric(site_scom_df$LAT) site_scom_df$LONG <- as.numeric(site_scom_df$LONG) # Add other numeric fields here (e.g., SCOM-related metrics) # Convert numeric fields in soil layer dataframe soil_layer_df$LAT <- as.numeric(soil_layer_df$LAT) soil_layer_df$LONG <- as.numeric(soil_layer_df$LONG) # Add soil-specific numeric fields here (e.g., depth, organic carbon)
Quick Notes to Adjust for Edge Cases
- If you have multiple consecutive empty lines, update the
empty_linesline toempty_lines <- which(dat == "" & dat != lag(dat, default = ""))to avoid duplicate splits. - Double-check the soil header rows (6 and 10) with a sample group—if the header structure varies, you might need a more flexible way to identify header lines (e.g., looking for keywords like "HORIZON").
- For large datasets, using
dplyr::bind_rows()instead ofrbind()will be faster and more efficient.
内容的提问来源于stack exchange,提问作者89_Simple
相关产品推荐
相关产品推荐

