You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言:在数据框中筛选含Combustion与Coal的行创建子集

Hey there! Let's sort out this subsetting task for your data frame (11717 obs, 15 vars) right away. You need rows where either both "Combustion" and "Coal" appear in the same cell, or they show up in different cells within the same row. Here are two straightforward approaches in R:

Base R Approach

First, let's create a helper function to check if a single cell contains both terms, then handle the two conditions separately before combining them:

# Helper function to check if a string has both "Combustion" and "Coal"
has_both <- function(x) {
  # Use ignore.case = TRUE if you want case-insensitive matching (e.g., "combustion" or "coal" counts)
  grepl("Combustion", x, ignore.case = FALSE) & grepl("Coal", x, ignore.case = FALSE)
}

# 1. Rows where at least one cell has both terms (same sentence/cell)
same_cell_rows <- apply(your_dataframe, 1, function(row) any(sapply(row, has_both), na.rm = TRUE))

# 2. Rows where at least one cell has "Combustion" AND at least one cell has "Coal" (different cells)
has_combustion <- apply(your_dataframe, 1, function(row) any(grepl("Combustion", row, ignore.case = FALSE), na.rm = TRUE))
has_coal <- apply(your_dataframe, 1, function(row) any(grepl("Coal", row, ignore.case = FALSE), na.rm = TRUE))
diff_cell_rows <- has_combustion & has_coal

# Combine both conditions and subset the data
final_subset <- your_dataframe[same_cell_rows | diff_cell_rows, ]

dplyr/tidyr Approach (More Readable)

If you prefer tidyverse syntax, this method is cleaner and easier to follow:

library(dplyr)
library(stringr)

final_subset <- your_dataframe %>%
  rowwise() %>%
  mutate(
    # Check for both terms in the same cell
    same_cell = any(str_detect(c_across(everything()), regex("Combustion", ignore_case = FALSE) & regex("Coal", ignore_case = FALSE)), na.rm = TRUE),
    # Check for each term in separate cells
    diff_cell = any(str_detect(c_across(everything()), "Combustion"), na.rm = TRUE) & 
                any(str_detect(c_across(everything()), "Coal"), na.rm = TRUE)
  ) %>%
  filter(same_cell | diff_cell) %>%
  select(-same_cell, -diff_cell) # Clean up helper columns

Quick Notes

  • Replace your_dataframe with the actual name of your data frame.
  • If you don't care about case sensitivity (e.g., "Combustion" vs "combustion"), change ignore.case = FALSE to TRUE in the grepl or regex calls.
  • The na.rm = TRUE argument ensures NA values don't break the checks (they'll be treated as "doesn't contain the term").

内容的提问来源于stack exchange,提问作者euclideans

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 04:19:27