You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用tidyverse修正NA替换及不匹配行移除的代码问题

问题描述

示例数据集

# Load the tidyverse package
library(tidyverse)

# Create the dataset
id <- 1:6
model <- c("0RB3211", NA, "0RB4191",
           NA, "0RB4033", NA)
UPC <- c("805289119081", "DK_0RB3447CP_RBCP  50", "8053672006360",
         "Green_Classic_G-15_Polar_1.67_PREM_SV", "805289044604",
         "DK_0RB2132CP_RBCP  55")
df <- tibble(id, model, UPC)

需求说明

当model列存在缺失值(NA)时:

  • 若对应UPC以DK开头,提取第一个下划线后的7位字符数字组合(如第二行提取"0RB3447"、最后一行提取"0RB2132")填充model;
  • 若UPC不以DK开头则删除整行(如第四行需删除)。

错误代码

# Manipulate the dataset
df_cleaned <- df %>%
  rowwise() %>%
  mutate(model = ifelse(is.na(model) & str_detect(UPC, "^DK"),
                        str_extract(UPC, "\\d{2}RB\\d{4}"),
                        model)) %>%
  ungroup() %>%
  filter(!(is.na(model) & str_detect(UPC, "[^0-9]")))

# Display the cleaned dataset
print(df_cleaned)

请问如何修改上述代码以得到正确结果?


修正方案

原代码问题分析

  1. 正则表达式错误:原代码用\\d{2}RB\\d{4}匹配目标字符串,但实际目标是0RB3447(1个数字+RB+4个数字,共7位),该正则无法正确捕获内容。
  2. 过滤逻辑错误:原过滤条件!(is.na(model) & str_detect(UPC, "[^0-9]"))逻辑混乱,无法准确筛选需要保留的行。

修正后代码

df_cleaned <- df %>%
  # 填充NA的model:仅当model为NA且UPC以DK开头时,提取DK_后7位
  mutate(model = case_when(
    is.na(model) & str_detect(UPC, "^DK") ~ str_extract(UPC, "(?<=DK_).{7}"),
    TRUE ~ model
  )) %>%
  # 删除model仍为NA的行(即原NA且UPC不以DK开头的行)
  filter(!is.na(model))

print(df_cleaned)

代码说明

  1. case_when替代ifelse:逻辑分支更清晰,明确处理特定条件下的填充操作。
  2. 正则表达式优化:(?<=DK_).{7}是正向预查,精准匹配DK_之后的任意7个字符;如果需要限定只能是字母数字组合,可改为(?<=DK_)[A-Z0-9]{7}。
  3. 过滤逻辑简化:直接过滤掉model仍为NA的行,完全符合需求中“删除UPC不以DK开头且model为NA的行”的要求。

内容的提问来源于stack exchange,提问作者Fox_Summer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 02:57:40