You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言中拆分含多值列失败问题求助

Hey there! Let's figure out how to split that exercise barriers column in your dataset—since the entries use | (with surrounding spaces) as a delimiter and have varying numbers of values, here are two practical approaches depending on what you need:

1. Split into columns based on value position

If you want each value in its own column ordered by where it appears in the original string (e.g., first value in column 1, second in column 2, etc.), use tidyr::separate_wider_delim—it handles varying lengths automatically:

First, let's recreate your sample data to test with:

# Sample dataset matching your example
exercise_data <- data.frame(
  barriers = c(
    "No transportation",
    "No transportation | Physically Unable | Don't know how",
    "Physically Unable | No transportation",
    "No Money | Don't know how | No Facilities | No free time"
  )
)

Now run the split:

library(tidyr)

# Split the barriers column into multiple position-based columns
split_by_position <- exercise_data %>%
  separate_wider_delim(
    cols = barriers,
    delim = " | ",  # Important: matches the exact delimiter (space + | + space)
    names_sep = "_",  # Names new columns like barriers_1, barriers_2, etc.
    too_few = "align_start",  # Shorter entries fill the first columns, leave others blank
    too_many = "error"  # Throws error if any entry has more values than the max found
  )

This will create columns barriers_1 through barriers_4 (since your longest entry has 4 values), with each column holding the corresponding reason from the original string.

2. Split into one-hot encoded columns (each unique reason as a column)

If you want a separate column for every unique barrier (marked 1 if present, 0 if not), combine separate_longer_delim and pivot_wider:

library(tidyr)
library(dplyr)

# Convert barriers to one-hot encoded columns
one_hot_barriers <- exercise_data %>%
  # First, split the string into individual rows per barrier
  separate_longer_delim(cols = barriers, delim = " | ") %>%
  # Add a marker that the barrier is present
  mutate(is_present = 1) %>%
  # Pivot back to wide format, filling missing values with 0
  pivot_wider(
    names_from = barriers,
    values_from = is_present,
    values_fill = 0
  )

This will give you columns like No transportation, Physically Unable, Don't know how, etc., where each row has a 1 if that barrier was listed, and 0 otherwise.

Just a quick note: double-check your delimiter—if some entries don't have spaces around the |, adjust the delim argument to just "|" and add a str_trim() step (from the stringr package) to clean up leftover whitespace!

内容的提问来源于stack exchange,提问作者Christopher Williams

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 08:17:01