从Python迁移到R:如何从字符串生成DataFrame
Got it, let's translate that Python functionality over to R step by step. Here's how you can turn that semi-colon separated key-value string into a proper DataFrame, just like your Python code does:
Step 1: Prepare and Clean the Input String
First, we'll start with your original string and strip out extra whitespace to avoid splitting issues:
# Original input string junk <- "total=7871MB;free=5711MB;used=2159MB;shared=0MB;buffers=304MB;cached=1059MB; free=71MB;total=5751MB;shared=3159MB;used=5MB;buffers=30MB;cached=1059MB; cached=1059MB;total=5751MB;shared=3159MB;used=5MB;buffers=30MB;free=109MB;" # Clean up: remove whitespace and split into individual key-value pairs clean_junk <- gsub("\\s", "", junk) key_value_pairs <- strsplit(clean_junk, ";")[[1]] # Filter out empty strings left from trailing/leading semicolons key_value_pairs <- key_value_pairs[key_value_pairs != ""]
Step 2: Extract Keys and Values
Next, we'll split each pair into key and value columns. You can use regex (matching your Python approach) or simple string splitting. Here's a tidyverse-friendly method:
library(stringr) library(tidyr) library(dplyr) # Use regex to capture key and value from each pair kv_df <- str_match(key_value_pairs, "(\\w+)=(.*)") %>% as.data.frame(stringsAsFactors = FALSE) %>% select(key = V2, value = V3)
If you prefer base R (no external packages), use this instead:
# Split each pair by "=" and convert to a data frame kv_list <- lapply(key_value_pairs, function(x) strsplit(x, "=")[[1]]) kv_df <- do.call(rbind, kv_list) %>% as.data.frame(stringsAsFactors = FALSE) colnames(kv_df) <- c("key", "value")
Step 3: Group Into Individual Records
Your string contains multiple full records (each starting with total). We'll create a grouping ID to cluster each complete record together:
# Create a group ID: increment every time we hit a "total" key kv_df$group <- cumsum(kv_df$key == "total")
Step 4: Reshape to Wide-Format DataFrame
Finally, we'll pivot the long-format key-value pairs into a wide-format DataFrame where each row is a complete record. We'll also handle duplicate keys (like the repeated cached in your string) by keeping the last occurrence:
# Tidyverse approach final_df <- kv_df %>% pivot_wider( names_from = key, values_from = value, values_fn = last # Resolve duplicate keys in a record ) %>% select(-group) # Remove the grouping column if not needed
Base R alternative using reshape:
# Base R reshape to wide format final_df <- reshape( kv_df, idvar = "group", timevar = "key", direction = "wide" ) # Clean up column names and remove group ID colnames(final_df) <- gsub("value\\.", "", colnames(final_df)) final_df <- final_df[, !colnames(final_df) %in% "group"]
The resulting final_df will have columns like total, free, used, etc., with each row representing one complete record from your original string.
内容的提问来源于stack exchange,提问作者Jan

