R语言:在数据框中筛选含Combustion与Coal的行创建子集
Hey there! Let's sort out this subsetting task for your data frame (11717 obs, 15 vars) right away. You need rows where either both "Combustion" and "Coal" appear in the same cell, or they show up in different cells within the same row. Here are two straightforward approaches in R:
Base R Approach
First, let's create a helper function to check if a single cell contains both terms, then handle the two conditions separately before combining them:
# Helper function to check if a string has both "Combustion" and "Coal" has_both <- function(x) { # Use ignore.case = TRUE if you want case-insensitive matching (e.g., "combustion" or "coal" counts) grepl("Combustion", x, ignore.case = FALSE) & grepl("Coal", x, ignore.case = FALSE) } # 1. Rows where at least one cell has both terms (same sentence/cell) same_cell_rows <- apply(your_dataframe, 1, function(row) any(sapply(row, has_both), na.rm = TRUE)) # 2. Rows where at least one cell has "Combustion" AND at least one cell has "Coal" (different cells) has_combustion <- apply(your_dataframe, 1, function(row) any(grepl("Combustion", row, ignore.case = FALSE), na.rm = TRUE)) has_coal <- apply(your_dataframe, 1, function(row) any(grepl("Coal", row, ignore.case = FALSE), na.rm = TRUE)) diff_cell_rows <- has_combustion & has_coal # Combine both conditions and subset the data final_subset <- your_dataframe[same_cell_rows | diff_cell_rows, ]
dplyr/tidyr Approach (More Readable)
If you prefer tidyverse syntax, this method is cleaner and easier to follow:
library(dplyr) library(stringr) final_subset <- your_dataframe %>% rowwise() %>% mutate( # Check for both terms in the same cell same_cell = any(str_detect(c_across(everything()), regex("Combustion", ignore_case = FALSE) & regex("Coal", ignore_case = FALSE)), na.rm = TRUE), # Check for each term in separate cells diff_cell = any(str_detect(c_across(everything()), "Combustion"), na.rm = TRUE) & any(str_detect(c_across(everything()), "Coal"), na.rm = TRUE) ) %>% filter(same_cell | diff_cell) %>% select(-same_cell, -diff_cell) # Clean up helper columns
Quick Notes
- Replace
your_dataframewith the actual name of your data frame. - If you don't care about case sensitivity (e.g., "Combustion" vs "combustion"), change
ignore.case = FALSEtoTRUEin thegreplorregexcalls. - The
na.rm = TRUEargument ensures NA values don't break the checks (they'll be treated as "doesn't contain the term").
内容的提问来源于stack exchange,提问作者euclideans
相关产品推荐
相关产品推荐

