You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言按列值差异拆分数据框时行数不符的问题排查与修复

数据拆分逻辑问题分析与修复

问题描述

需要把数据集拆成三个数据框:

  • 仅在comp01_name列出现的竞争者对应的行
  • 仅在comp02_name列出现的竞争者对应的行
  • 在两列都出现的竞争者对应的行

但用setdiff统计出的唯一竞争者数量,和实际生成的数据框行数对不上:

  • 预期master.treeDQ2.1是2037行,实际跑出2141行
  • 预期master.treeDQ2.2是1476行,实际跑出3750行

原始代码如下:

master.treeDQ2 = read.csv('https://raw.githubusercontent.com/bandcar/Examples/main/master.treeDQ.edited_draft6.csv')

# check for competitors in comp02 that are not in comp01
length(setdiff(master.treeDQ2$comp02_name, master.treeDQ2$comp01_name)) #1476 ppl are in comp02 that are not in comp01

# check for competitors in comp01 that are not in comp02
length(setdiff(master.treeDQ2$comp01_name, master.treeDQ2$comp02_name))
# 2037 are in comp01, but not in comp02

# subset competitors who only appear in comp01 but not comp02
master.treeDQ2.1 = master.treeDQ2[!(master.treeDQ2$comp02_name %in% master.treeDQ2$comp01_name),]

# subset competitors who only appear in comp02 but not comp01
master.treeDQ2.2 = master.treeDQ2[!(master.treeDQ2$comp01_name %in% master.treeDQ2$comp02_name),]

# subset for competitors present in both columns
master.treeDQboth = master.treeDQ2[master.treeDQ2$comp01_name %in% master.treeDQ2$comp02_name,]

问题原因

  1. 筛选逻辑完全跑偏:
    • 你写的master.treeDQ2.1筛选条件!(master.treeDQ2$comp02_name %in% master.treeDQ2$comp01_name),实际是挑「当前行的comp02_name没在comp01_name列出现过」的行,而不是挑「只在comp01列出现的竞争者对应的所有行」。这会把很多不属于目标类别的行误选进来。
    • master.treeDQ2.2的条件更是搞反了,用comp01_name的存在性来判断仅在comp02出现的行,逻辑完全不对。
  2. 混淆了「唯一竞争者数量」和「数据行数」:
    setdiff统计的是符合条件的唯一竞争者个数,但数据集中同一个竞争者可能对应多行记录,所以直接按行筛选得到的行数肯定和唯一值数量对不上。

修复方案

先把三类竞争者的唯一名单提出来,再用名单去筛选对应的行:

master.treeDQ2 = read.csv('https://raw.githubusercontent.com/bandcar/Examples/main/master.treeDQ.edited_draft6.csv')

# 提取两列的唯一竞争者名单
comp1_unique <- unique(master.treeDQ2$comp01_name)
comp2_unique <- unique(master.treeDQ2$comp02_name)

# 划分三类竞争者的唯一名单
only_comp1 <- setdiff(comp1_unique, comp2_unique)  # 只在comp01出现的竞争者
only_comp2 <- setdiff(comp2_unique, comp1_unique)  # 只在comp02出现的竞争者
both_comp <- intersect(comp1_unique, comp2_unique) # 在两列都出现的竞争者

# 筛选对应的数据框
# 1. 仅在comp01出现的竞争者的所有行
master.treeDQ2.1 <- master.treeDQ2[master.treeDQ2$comp01_name %in% only_comp1, ]

# 2. 仅在comp02出现的竞争者的所有行
master.treeDQ2.2 <- master.treeDQ2[master.treeDQ2$comp02_name %in% only_comp2, ]

# 3. 在两列都出现的竞争者的所有行
master.treeDQboth <- master.treeDQ2[master.treeDQ2$comp01_name %in% both_comp | master.treeDQ2$comp02_name %in% both_comp, ]

验证方法

可以用下面的代码验证结果是否合理:

# 查看三类数据框的行数
nrow(master.treeDQ2.1)
nrow(master.treeDQ2.2)
nrow(master.treeDQboth)

# 验证三类行的总数是否等于原数据行数(确保没有遗漏或重复)
nrow(master.treeDQ2.1) + nrow(master.treeDQ2.2) + nrow(master.treeDQboth) == nrow(master.treeDQ2)

内容的提问来源于stack exchange,提问作者bandcar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 12:10:45