R语言数据框列替换:将生物学相关词汇统一替换为‘biology’
Hey there! No worries at all—we all start somewhere with R. Let's break down how to solve this problem step by step.
Step 1: Recreate Your Sample Data
First, let's make sure we're working with the same data structure you described. Your subject column is a list of character vectors, so we can build the data frame like this:
# Create the sample data frame df <- data.frame( title = c("A", "B", "C"), subject = I(list( c("health sciences", "life sciences"), c("biochemistry", "medicine", "life sciences"), c("physics and astronomy", "mathematics") )) )
Step 2: Define Biology-Related Terms
Next, let's list out all the terms you want to replace with biology. You can add or remove terms here based on your specific needs:
# Define terms to replace with "biology" biology_terms <- c("health sciences", "life sciences", "biochemistry", "medicine")
Step 3: Replace Terms Using Base R
If you prefer using base R (no extra packages needed), you can use lapply() to loop through each element in the subject list and replace matching terms:
# Replace matching terms with "biology" df$subject <- lapply(df$subject, function(x) { # Replace any term in biology_terms with "biology" x[x %in% biology_terms] <- "biology" return(x) }) # View the processed data print(df)
Output:
title subject 1 A biology, biology 2 B biology, biology, biology 3 C physics and astronomy, mathematics
Step 4: Alternative Using dplyr + purrr
If you're using the tidyverse, you can use mutate() and map() for a more readable pipeline:
# Load required packages library(dplyr) library(purrr) # Process the data df_processed <- df %>% mutate(subject = map(subject, ~ ifelse(.x %in% biology_terms, "biology", .x))) # View the result print(df_processed)
This will give you the same output as the base R method.
Quick Note for Edge Cases
If your subject column is stored as strings (e.g., "c('health sciences', 'life sciences')" instead of a list), you'll need to convert it to a list first. You can do that with:
df$subject <- lapply(df$subject, function(x) eval(parse(text = x)))
Note: Be cautious with eval(parse())—only use it if you trust the source of your data, as it can execute arbitrary code.
Let me know if you need help adjusting this to fit your exact data!
内容的提问来源于stack exchange,提问作者Ismail Jan

