技术问询:如何用gather拉长列表及用tidyverse/base-r处理重复字段
Hey there! Let's break down your two questions clearly, with practical examples for each:
gather() to "lengthen" your data (list/data frame) First off, gather() is part of the tidyr package in the tidyverse. While it's been superseded by the more flexible pivot_longer() these days, it still works great for turning wide-format data into long-format (what you're calling "lengthening").
Example with a wide data frame
Say you have a wide data frame where each judge's score is a separate column:
library(tidyr) # Sample wide data wide_scores <- data.frame( participant_id = 1:3, Judge_poster_score_1 = c(8, 9, 7), Judge_poster_score_2 = c(7, 8, 9), Judge_poster_score_3 = c(9, 7, 8) )
To lengthen this, use gather() to collapse the judge score columns into two new columns: one for the judge identifier, and one for the score:
long_scores <- wide_scores %>% gather( key = "judge_label", # Name of the new column for original column names value = "poster_score", # Name of the new column for the values starts_with("Judge_poster_score") # Which columns to lengthen )
If you're working with a list (instead of a data frame), first convert it to a data frame, then apply gather() the same way:
score_list <- list( participant_id = 1:3, Judge_poster_score_1 = c(8,9,7), Judge_poster_score_2 = c(7,8,9) ) # Convert list to data frame, then lengthen long_from_list <- score_list %>% as.data.frame() %>% gather(key = "judge_label", value = "poster_score", -participant_id)
Since you already have a data.table solution, let's cover tidyverse and base R approaches for two common scenarios: extracting from text strings, and selecting matching columns in a data frame.
Tidyverse approach (stringr + dplyr)
For text strings with repeated phrases
If you have a vector of text where "Judge poster score" appears multiple times, use stringr::str_extract_all() to pull out every instance:
library(stringr) library(dplyr) library(tidyr) sample_text <- c( "Judge poster score: 8, Judge poster score: 9", "Review: Judge poster score: 7, Notes: Judge poster score: 10" ) # Extract all instances and convert to a tidy data frame extracted_phrases <- tibble(raw_text = sample_text) %>% mutate(matching_phrases = str_extract_all(raw_text, "Judge poster score")) %>% unnest(matching_phrases)
For selecting columns with matching names
If your data frame has columns named like Judge poster score 1, Judge poster score 2, use dplyr::select() with contains() to grab them all:
# Sample data frame with matching columns score_df <- data.frame( id = 1:3, `Judge poster score 1` = c(8,9,7), `Judge poster score 2` = c(7,8,9), other_data = c("A", "B", "C") ) # Select only the "Judge poster score" columns target_columns <- score_df %>% select(contains("Judge poster score"))
Base R approach
For text strings with repeated phrases
Use gregexpr() and regmatches() to extract all instances of the phrase:
sample_text <- c( "Judge poster score: 8, Judge poster score: 9", "Review: Judge poster score: 7, Notes: Judge poster score: 10" ) # Extract all matches into a list extracted_list <- lapply(sample_text, function(x) { regmatches(x, gregexpr("Judge poster score", x))[[1]] }) # Convert to a data frame if needed extracted_df <- data.frame( raw_text = rep(sample_text, sapply(extracted_list, length)), matching_phrases = unlist(extracted_list) )
For selecting columns with matching names
Use grep() to filter column names:
# Using the same score_df from above target_columns_base <- score_df[, grep("Judge poster score", colnames(score_df))]
内容的提问来源于stack exchange,提问作者Doug Federman

