You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

检测R数据框两列相同组合并统计Relation各水平观测数

Hey there! Let's work through this problem together. The core challenge here is making sure reverse pairs (like ID24 & ID1 vs. ID1 & ID24) get recognized as the same combination—once we fix that, counting up the Relation levels is a breeze.

Step 1: Standardize the Individual Pairs

First, we need to create a consistent identifier for each pair, regardless of their original order. We can use pmin() and pmax() to sort the two IDs in each row, so every pair gets the same "standard" order. Here's how to do that:

# Assuming your data frame is named df
df <- df %>%
  mutate(
    # Sort the IDs to create a consistent pair
    Pair1 = pmin(OffspringID1, OffspringID2),
    Pair2 = pmax(OffspringID1, OffspringID2),
    # Optional: Combine into a single string for easier reference
    Combination = paste(Pair1, Pair2, sep = "-")
  )

For example, your rows 1 (ID24 & ID1) and 5 (ID1 & ID24) will both end up with Pair1 = ID1 and Pair2 = ID24, so they're grouped together correctly.

Step 2: Count Observations by Pair & Relation

Now we can count how many times each Relation level appears for each standardized pair. Here are two common methods:

Method 1: Using dplyr (Intuitive & Flexible)

If you use the dplyr package, this becomes super straightforward with grouping and summarizing:

library(dplyr)

count_result <- df %>%
  group_by(Pair1, Pair2, Relation) %>%
  summarize(Count = n(), .groups = "drop")

# Or if you prefer using the combined Combination column:
count_result <- df %>%
  group_by(Combination, Relation) %>%
  summarize(Count = n(), .groups = "drop")

Method 2: Using aggregate() (As You Initially Considered)

If you want to stick with base R's aggregate() function, you can use the standardized pairs as grouping variables. We'll count using the Replicate column (since each row is one observation, length() will give the count):

count_result <- aggregate(
  Replicate ~ Pair1 + Pair2 + Relation,
  data = df,
  FUN = length
)
# Rename the column for clarity
names(count_result)[4] <- "Count"

Example Output Snippet

For your sample data, the result would include a row for the ID1-ID24 pair like this:

Pair1Pair2RelationCount
ID1ID24PO1
ID1ID24HS1

This gives you a clear breakdown of how many times each Relation level occurs for every unique individual pair—including reverse pairs treated as one.

内容的提问来源于stack exchange,提问作者Chrys

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 04:30:19