You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言求助:将不等长字符数组转换为DataFrame

Hey there! As a new R user, converting that character vector into a structured DataFrame might feel tricky at first, but let's walk through a straightforward solution using the tidyverse package—it's beginner-friendly and gets the job done neatly.

Step 1: Define your input data

First, let's write out your character vector clearly so you can copy-paste it directly into R:

text_data <- c(
  "Colour: Gold|Style: Without Offers",
  "Colour: Gold|Style: Without Offers|Verified Purchase",
  "Colour: Gold|Style: With Offers|Verified Purchase",
  "Colour: Gold|Style: Without Offers|Verified Purchase",
  "Colour: Black|Verified Purchase",
  "Colour: Gold|Style: Without Offers"
)

Step 2: Install & load the tidyverse package

If you haven't installed the tidyverse toolkit yet (it's a must-have for data work in R), run this first:

install.packages("tidyverse")

Then load it into your session:

library(tidyverse)

Step 3: Process the data into your desired DataFrame

Use these lines to split the text, clean up column names, and handle missing values:

# Convert the character vector to a tibble (a modern, easy-to-use DataFrame)
df <- tibble(text = text_data)

# Split the text into columns, rename fields, and replace empty values with NA
result <- df %>%
  separate_wider_delim(
    text,
    delim = "|",
    names_sep = ": ",  # Split each key-value pair at the ": " separator
    names_repair = ~ str_replace(., "Verified Purchase", "PurchaseType")  # Match your desired column name
  ) %>%
  mutate(across(everything(), ~ ifelse(. == "", NA, .)))  # Turn empty cells into R's standard missing value marker

# View the final result
result

What you'll get

When you run the code, you'll see exactly the structured output you wanted:

# A tibble: 6 × 3
  Colour Style           PurchaseType      
  <chr>  <chr>           <chr>             
1 Gold   Without Offers  NA                
2 Gold   Without Offers  Verified Purchase
3 Gold   With Offers     Verified Purchase
4 Gold   Without Offers  Verified Purchase
5 Black  NA              Verified Purchase
6 Gold   Without Offers  NA                

Quick breakdown of the code

  • separate_wider_delim: Splits each string by |, then uses : to split each segment into column names and values.
  • names_repair: Renames the "Verified Purchase" field to "PurchaseType" to match your expected output.
  • mutate(across(...)): Replaces empty cells (from missing fields like Style) with NA, which is R's standard way to represent missing data.

Alternative: Base R solution (no external packages)

If you prefer to stick to base R without installing extra packages, here's a way to do it:

# Split each string into key-value pairs
split_list <- strsplit(text_data, "\\|")

# Process each entry to map to target columns
processed <- lapply(split_list, function(x) {
  key_val <- strsplit(x, ": ")
  keys <- sapply(key_val, `[`, 1)
  vals <- sapply(key_val, `[`, 2)
  
  data.frame(
    Colour = ifelse("Colour" %in% keys, vals[keys == "Colour"], NA),
    Style = ifelse("Style" %in% keys, vals[keys == "Style"], NA),
    PurchaseType = ifelse("Verified Purchase" %in% keys, "Verified Purchase", NA),
    stringsAsFactors = FALSE
  )
})

# Combine all entries into a single DataFrame
result_base <- do.call(rbind, processed)
result_base

This base R method achieves the same result but requires more manual handling of key-value pairs. The tidyverse approach is generally more readable for beginners!

内容的提问来源于stack exchange,提问作者Sana Ansari

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:03:10