R语言:按列名选列匹配NA及客户营收分类方案问询
Hey there! Let's work through your two R tasks step by step, with practical, runnable code examples for each.
Selecting columns by name and identifying NA values in those columns is a common task—here are two straightforward approaches using base R and the tidyverse:
方法1:Base R 实现
First, let's create a sample data frame to work with, then select columns and flag rows with NA values:
# 构造示例数据集 df <- data.frame( col1 = c(1, NA, 3, 4), col2 = c("a", "b", "c", "d"), col3 = c(NA, 2, NA, 4), col4 = c(TRUE, FALSE, TRUE, FALSE), col5 = c(5, 6, NA, 8) ) # 按列名选择目标列 selected_cols <- df[, c("col1", "col3", "col5")] # 筛选出目标列中**存在至少一个NA**的行 rows_with_any_na <- df[rowSums(is.na(selected_cols)) > 0, ] # 筛选出目标列中**所有值都是NA**的行 rows_with_all_na <- df[rowSums(is.na(selected_cols)) == ncol(selected_cols), ]
方法2:Tidyverse(dplyr)实现
Using dplyr makes the code more readable, especially for chained operations:
library(dplyr) # 选择目标列并筛选含至少一个NA的行 df %>% select(col1, col3, col5) %>% filter(if_any(everything(), is.na)) # 选择目标列并筛选所有值都是NA的行 df %>% select(col1, col3, col5) %>% filter(if_all(everything(), is.na))
Let's start by building a minimal runnable dataset that matches your scenario, then categorize each client:
# 构造示例营收数据(模拟多年多客户结构) revenue_data <- data.frame( name = c("name1", "name1", "name2", "name2", "name3", "name4", "name4", "name5"), year = c(2017, 2018, 2017, 2018, 2017, 2018, 2018, 2018), revenue = c(100, 150, 200, 220, 300, 50, 70, 90) )
Now we'll categorize each client based on their presence in 2017 and 2018 using dplyr:
library(dplyr) # 第一步:按客户分组,标记是否在2017/2018年有营收 client_status <- revenue_data %>% group_by(name) %>% summarize( has_2017 = any(year == 2017), # 判断该客户2017年是否有营收 has_2018 = any(year == 2018) # 判断该客户2018年是否有营收 ) %>% # 第二步:根据标记结果分类 mutate( category = case_when( has_2018 & !has_2017 ~ "new", # 2018有营收、2017无 → 新增 has_2017 & has_2018 ~ "recurrent", # 两年都有营收 → 留存 has_2017 & !has_2018 ~ "lost", # 2017有营收、2018无 → 流失 TRUE ~ "unknown" # 处理特殊情况(如两年都无营收) ) ) # 第三步:将分类结果合并回原数据集(如果需要保留原始营收明细) final_data <- revenue_data %>% left_join(client_status, by = "name") # 查看最终结果 print(final_data)
This code scales seamlessly to large datasets with multiple years and clients—since we group by each unique client first, it will correctly evaluate every client's status regardless of how many monthly entries they have.
内容的提问来源于stack exchange,提问作者Philippe

