You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

技术问询:如何用gather拉长列表及用tidyverse/base-r处理重复字段

Hey there! Let's break down your two questions clearly, with practical examples for each:

1. Using gather() to "lengthen" your data (list/data frame)

First off, gather() is part of the tidyr package in the tidyverse. While it's been superseded by the more flexible pivot_longer() these days, it still works great for turning wide-format data into long-format (what you're calling "lengthening").

Example with a wide data frame

Say you have a wide data frame where each judge's score is a separate column:

library(tidyr)

# Sample wide data
wide_scores <- data.frame(
  participant_id = 1:3,
  Judge_poster_score_1 = c(8, 9, 7),
  Judge_poster_score_2 = c(7, 8, 9),
  Judge_poster_score_3 = c(9, 7, 8)
)

To lengthen this, use gather() to collapse the judge score columns into two new columns: one for the judge identifier, and one for the score:

long_scores <- wide_scores %>%
  gather(
    key = "judge_label",  # Name of the new column for original column names
    value = "poster_score",  # Name of the new column for the values
    starts_with("Judge_poster_score")  # Which columns to lengthen
  )

If you're working with a list (instead of a data frame), first convert it to a data frame, then apply gather() the same way:

score_list <- list(
  participant_id = 1:3,
  Judge_poster_score_1 = c(8,9,7),
  Judge_poster_score_2 = c(7,8,9)
)

# Convert list to data frame, then lengthen
long_from_list <- score_list %>%
  as.data.frame() %>%
  gather(key = "judge_label", value = "poster_score", -participant_id)
2. Extracting repeated "Judge poster score" with tidyverse or base R

Since you already have a data.table solution, let's cover tidyverse and base R approaches for two common scenarios: extracting from text strings, and selecting matching columns in a data frame.

Tidyverse approach (stringr + dplyr)

For text strings with repeated phrases

If you have a vector of text where "Judge poster score" appears multiple times, use stringr::str_extract_all() to pull out every instance:

library(stringr)
library(dplyr)
library(tidyr)

sample_text <- c(
  "Judge poster score: 8, Judge poster score: 9",
  "Review: Judge poster score: 7, Notes: Judge poster score: 10"
)

# Extract all instances and convert to a tidy data frame
extracted_phrases <- tibble(raw_text = sample_text) %>%
  mutate(matching_phrases = str_extract_all(raw_text, "Judge poster score")) %>%
  unnest(matching_phrases)

For selecting columns with matching names

If your data frame has columns named like Judge poster score 1, Judge poster score 2, use dplyr::select() with contains() to grab them all:

# Sample data frame with matching columns
score_df <- data.frame(
  id = 1:3,
  `Judge poster score 1` = c(8,9,7),
  `Judge poster score 2` = c(7,8,9),
  other_data = c("A", "B", "C")
)

# Select only the "Judge poster score" columns
target_columns <- score_df %>%
  select(contains("Judge poster score"))

Base R approach

For text strings with repeated phrases

Use gregexpr() and regmatches() to extract all instances of the phrase:

sample_text <- c(
  "Judge poster score: 8, Judge poster score: 9",
  "Review: Judge poster score: 7, Notes: Judge poster score: 10"
)

# Extract all matches into a list
extracted_list <- lapply(sample_text, function(x) {
  regmatches(x, gregexpr("Judge poster score", x))[[1]]
})

# Convert to a data frame if needed
extracted_df <- data.frame(
  raw_text = rep(sample_text, sapply(extracted_list, length)),
  matching_phrases = unlist(extracted_list)
)

For selecting columns with matching names

Use grep() to filter column names:

# Using the same score_df from above
target_columns_base <- score_df[, grep("Judge poster score", colnames(score_df))]

内容的提问来源于stack exchange,提问作者Doug Federman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 08:05:33