You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言求助:基于共享产品的分组计数实现

Hey there! Let's walk through how to solve this step by step since you're new to R—no worries, it's straightforward once you break it down.

First, let's start by recreating your dataset in R so we can work with it:

# Create your example dataset as a data frame
df <- data.frame(
  Comp1 = c("A", "A", "A", "C", "C", "X", "X", "X"),
  Product = c("P1", "P2", "P3", "P4", "P2", "P1", "P2", "P3"),
  Comp2 = c("B", "B", "B", "D", "D", "Y", "Y", "Y")
)

Next, we'll use the dplyr and tidyr packages (super common for data manipulation in R) to process the data. If you haven't installed these yet, run this first:

install.packages(c("dplyr", "tidyr"))

Now, here's the full code to get your desired output, with comments explaining each step:

library(dplyr)
library(tidyr)

# Step 1: Create a consistent identifier for each company pair (so A-B and B-A count as the same pair)
# Step 2: Group by each company pair, collect their shared products, sort them, and combine into a comma-separated string
# Step 3: Count how many times each product bundle appears
result <- df %>%
  rowwise() %>%
  # Sort Comp1 and Comp2 to make a unique pair ID (avoids issues with reverse pairs)
  mutate(company_pair = paste(sort(c(Comp1, Comp2)), collapse = "-")) %>%
  ungroup() %>%
  # Group by the company pair to get their shared products
  group_by(company_pair) %>%
  # Sort products then join with commas (so P2,P1 becomes P1,P2 for consistency)
  summarise(product_bundle = paste(sort(Product), collapse = ",")) %>%
  ungroup() %>%
  # Count occurrences of each product bundle
  count(product_bundle, name = "Count") %>%
  # Rename columns to match your desired output format
  rename(`Product Bundle` = product_bundle)

# View the final result
print(result)

When you run this, you'll get exactly the output you wanted:

# A tibble: 2 × 2
  `Product Bundle` Count
  <chr>            <int>
1 P1,P2,P3             2
2 P2,P4                1

Let me quickly break down why each part matters:

  • Sorting company pairs: This ensures that if you ever have a pair like B-A instead of A-B, they'll still be grouped together (we don't want to treat them as separate pairs when they share the same products).
  • Sorting products: This makes sure that product bundles like P2,P1,P3 and P1,P2,P3 are counted as the same bundle (since they're the same set of products, just ordered differently).
  • Using count: This simplifies tallying how many unique company pairs have each product bundle.

If you have any questions about any part of this (like what rowwise() does, or why we use summarise), feel free to ask!

内容的提问来源于stack exchange,提问作者Mili

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:09:44