R语言求助:基于共享产品的分组计数实现
Hey there! Let's walk through how to solve this step by step since you're new to R—no worries, it's straightforward once you break it down.
First, let's start by recreating your dataset in R so we can work with it:
# Create your example dataset as a data frame df <- data.frame( Comp1 = c("A", "A", "A", "C", "C", "X", "X", "X"), Product = c("P1", "P2", "P3", "P4", "P2", "P1", "P2", "P3"), Comp2 = c("B", "B", "B", "D", "D", "Y", "Y", "Y") )
Next, we'll use the dplyr and tidyr packages (super common for data manipulation in R) to process the data. If you haven't installed these yet, run this first:
install.packages(c("dplyr", "tidyr"))
Now, here's the full code to get your desired output, with comments explaining each step:
library(dplyr) library(tidyr) # Step 1: Create a consistent identifier for each company pair (so A-B and B-A count as the same pair) # Step 2: Group by each company pair, collect their shared products, sort them, and combine into a comma-separated string # Step 3: Count how many times each product bundle appears result <- df %>% rowwise() %>% # Sort Comp1 and Comp2 to make a unique pair ID (avoids issues with reverse pairs) mutate(company_pair = paste(sort(c(Comp1, Comp2)), collapse = "-")) %>% ungroup() %>% # Group by the company pair to get their shared products group_by(company_pair) %>% # Sort products then join with commas (so P2,P1 becomes P1,P2 for consistency) summarise(product_bundle = paste(sort(Product), collapse = ",")) %>% ungroup() %>% # Count occurrences of each product bundle count(product_bundle, name = "Count") %>% # Rename columns to match your desired output format rename(`Product Bundle` = product_bundle) # View the final result print(result)
When you run this, you'll get exactly the output you wanted:
# A tibble: 2 × 2 `Product Bundle` Count <chr> <int> 1 P1,P2,P3 2 2 P2,P4 1
Let me quickly break down why each part matters:
- Sorting company pairs: This ensures that if you ever have a pair like B-A instead of A-B, they'll still be grouped together (we don't want to treat them as separate pairs when they share the same products).
- Sorting products: This makes sure that product bundles like
P2,P1,P3andP1,P2,P3are counted as the same bundle (since they're the same set of products, just ordered differently). - Using
count: This simplifies tallying how many unique company pairs have each product bundle.
If you have any questions about any part of this (like what rowwise() does, or why we use summarise), feel free to ask!
内容的提问来源于stack exchange,提问作者Mili

