You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何筛选交易修正行?基于R处理UCI在线零售数据集

Got it, let's figure out why your current code isn't working and fix this for the Online Retail dataset. Those correction records (with negative quantities and different invoice codes) aren't exact duplicates, so duplicated() won't catch them—here's how to target them properly:

First, let's fix your code's issues

Your original code has two key problems:

  1. Syntax error: You're missing a comma after data3$UnitPrice, and you shouldn't pass column values (data3$StockCode) to the column index—you need to use column names instead.
  2. Logical mismatch: duplicated() looks for exact duplicate rows, but correction records have unique invoice codes and negative quantities, so they aren't exact duplicates. We need to target matching transaction entities (same customer, product, price, country) that have both positive and negative quantities.

Solution 1: Filter all transactions with corresponding correction records

If you want to extract every transaction (original and corrected) where a product-customer pair has both positive and negative quantities, use one of these approaches:

Base R Version

# First, clean out rows with missing CustomerID (many transactions lack this)
data_clean <- data3[!is.na(data3$CustomerID), ]

# Create a unique group key for matching transaction entities
data_clean$group_key <- with(data_clean, paste(StockCode, Description, UnitPrice, CustomerID, Country, sep = "|"))

# Identify groups that have both positive and negative quantities
has_correction <- aggregate(Quantity ~ group_key, data_clean, 
                            FUN = function(x) length(unique(sign(x))) > 1)

# Pull all rows from those groups
Transactions <- data_clean[data_clean$group_key %in% has_correction$group_key[has_correction$Quantity], ]

dplyr Version (cleaner, more readable)

library(dplyr)

Transactions <- data3 %>%
  filter(!is.na(CustomerID)) %>%
  mutate(group_key = paste(StockCode, Description, UnitPrice, CustomerID, Country, sep = "|")) %>%
  group_by(group_key) %>%
  filter(n_distinct(sign(Quantity)) > 1) %>%
  ungroup() %>%
  select(-group_key) # Optional: remove the group key column if you don't need it

Solution 2: Remove correction records (keep only valid positive transactions)

If your goal is to exclude cancellation/correction records, the Online Retail dataset typically marks these with an invoice code starting with "C". You can filter those out directly:

# Filter out cancellation invoices and rows with missing CustomerID
Transactions <- data3 %>%
  filter(!grepl("^C", InvoiceNo)) %>%
  filter(!is.na(CustomerID))

If some corrections don't have the "C" prefix, use the group key approach to remove negative quantity rows that have a matching positive counterpart:

Transactions <- data3 %>%
  filter(!is.na(CustomerID)) %>%
  mutate(group_key = paste(StockCode, Description, UnitPrice, CustomerID, Country, sep = "|")) %>%
  group_by(group_key) %>%
  filter(Quantity > 0 | sum(Quantity) < 0) %>% # Keep positive rows, or full groups if total is negative (full cancellation)
  ungroup() %>%
  select(-group_key)

内容的提问来源于stack exchange,提问作者Sergi F.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:26:41