You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R语言中识别两个数据集间的家庭成员数量不匹配问题

Detecting Family Member Count Mismatches in R

Got it, let's tackle this mismatch detection problem step by step. Here's how you can easily spot families where the recorded member count in Dataset1 doesn't line up with the actual number of entries in Dataset2:

Step 1: Recreate your datasets in R

First, let's build the two datasets exactly as you described:

# Dataset1 with recorded family member counts
Dataset1 <- data.frame(
  family_id = c(1, 2, 3),
  house_id = c(1052, 5042, 1111),
  number_family_member = c(2, 3, 2)
)

# Dataset2 with individual family member details
Dataset2 <- data.frame(
  family_id = c(1, 1, 2, 2, 3, 3, 3),
  house_id = c(1052, 1052, 5042, 5042, 1111, 1111, 1111),
  age = c(24, 25, 23, 20, 1, 20, 21),
  gender = c("male", "female", "male", "female", "male", "female", "female")
)

Step 2: Use dplyr for a clean, readable solution

The dplyr package makes this kind of data manipulation straightforward. We'll first calculate the actual number of members per family from Dataset2, then compare it to Dataset1's recorded counts:

library(dplyr)

# Calculate actual member counts grouped by family and house
actual_member_counts <- Dataset2 %>%
  group_by(family_id, house_id) %>%
  summarise(actual_count = n(), .groups = "drop")

# Merge with Dataset1 and filter for mismatches
mismatched_families <- Dataset1 %>%
  left_join(actual_member_counts, by = c("family_id", "house_id")) %>%
  filter(number_family_member != actual_count)

# View the result
mismatched_families

Running this code will output the families with mismatched counts:

family_id house_id number_family_member actual_count
1         2     5042                    3            2
2         3     1111                    2            3

Step 3: Base R alternative (no external packages)

If you prefer not to use add-on packages, here's a base R approach that achieves the same result:

# Calculate actual member counts using aggregate()
actual_counts_base <- aggregate(age ~ family_id + house_id, data = Dataset2, FUN = length)
names(actual_counts_base)[3] <- "actual_count"

# Merge datasets and filter for mismatches
mismatched_base <- merge(Dataset1, actual_counts_base, by = c("family_id", "house_id"))
mismatched_base <- mismatched_base[mismatched_base$number_family_member != mismatched_base$actual_count, ]

# View the result
mismatched_base

Both methods work equally well—choose whichever fits your workflow best.

内容的提问来源于stack exchange,提问作者Bikram Adhitya Adhikari

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 08:24:20