You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

从Python迁移到R:如何从字符串生成DataFrame

Convert Messy Key-Value String to DataFrame in R

Got it, let's translate that Python functionality over to R step by step. Here's how you can turn that semi-colon separated key-value string into a proper DataFrame, just like your Python code does:

Step 1: Prepare and Clean the Input String

First, we'll start with your original string and strip out extra whitespace to avoid splitting issues:

# Original input string
junk <- "total=7871MB;free=5711MB;used=2159MB;shared=0MB;buffers=304MB;cached=1059MB; free=71MB;total=5751MB;shared=3159MB;used=5MB;buffers=30MB;cached=1059MB; cached=1059MB;total=5751MB;shared=3159MB;used=5MB;buffers=30MB;free=109MB;"

# Clean up: remove whitespace and split into individual key-value pairs
clean_junk <- gsub("\\s", "", junk)
key_value_pairs <- strsplit(clean_junk, ";")[[1]]
# Filter out empty strings left from trailing/leading semicolons
key_value_pairs <- key_value_pairs[key_value_pairs != ""]

Step 2: Extract Keys and Values

Next, we'll split each pair into key and value columns. You can use regex (matching your Python approach) or simple string splitting. Here's a tidyverse-friendly method:

library(stringr)
library(tidyr)
library(dplyr)

# Use regex to capture key and value from each pair
kv_df <- str_match(key_value_pairs, "(\\w+)=(.*)") %>%
  as.data.frame(stringsAsFactors = FALSE) %>%
  select(key = V2, value = V3)

If you prefer base R (no external packages), use this instead:

# Split each pair by "=" and convert to a data frame
kv_list <- lapply(key_value_pairs, function(x) strsplit(x, "=")[[1]])
kv_df <- do.call(rbind, kv_list) %>% as.data.frame(stringsAsFactors = FALSE)
colnames(kv_df) <- c("key", "value")

Step 3: Group Into Individual Records

Your string contains multiple full records (each starting with total). We'll create a grouping ID to cluster each complete record together:

# Create a group ID: increment every time we hit a "total" key
kv_df$group <- cumsum(kv_df$key == "total")

Step 4: Reshape to Wide-Format DataFrame

Finally, we'll pivot the long-format key-value pairs into a wide-format DataFrame where each row is a complete record. We'll also handle duplicate keys (like the repeated cached in your string) by keeping the last occurrence:

# Tidyverse approach
final_df <- kv_df %>%
  pivot_wider(
    names_from = key,
    values_from = value,
    values_fn = last  # Resolve duplicate keys in a record
  ) %>%
  select(-group)  # Remove the grouping column if not needed

Base R alternative using reshape:

# Base R reshape to wide format
final_df <- reshape(
  kv_df,
  idvar = "group",
  timevar = "key",
  direction = "wide"
)
# Clean up column names and remove group ID
colnames(final_df) <- gsub("value\\.", "", colnames(final_df))
final_df <- final_df[, !colnames(final_df) %in% "group"]

The resulting final_df will have columns like total, free, used, etc., with each row representing one complete record from your original string.

内容的提问来源于stack exchange,提问作者Jan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 11:25:45