You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于两列回文行筛选:保留value最小值行的R/Python实现求助

问题描述

给定如下R数据框:

Data_Frame <- data.frame (
  A = c("a", "c", "b", "e", "g", "d", "f", "h"),
  B = c("b", "d", "a", "f", "h", "c", "e", "g"),
  value = c("0.3", "0.2", "0.1", "0.1", "0.5", "0.7", "0.8", "0.1"),
  effect = c("123", "345", "123", "444", "123", "345", "444", "123")
)

需完成以下操作:

  • 找出A列与B列呈回文关系(行i的A=行j的B,行i的B=行j的A)且effect列值相等的行对
  • 从每对回文行中保留value列数值最小的行
  • 注意:levels(Data_Frame$A)与levels(Data_Frame$B)不相等,as.character()无法解决该问题

期望输出的数据框:

Data_Frame <- data.frame (
  A = c("c", "b", "e", "h"),
  B = c("d", "a", "f", "g"),
  value = c("0.2", "0.1", "0.1", "0.1"),
  effect = c("345", "123", "444", "123")
)

R 实现方案

步骤1:统一因子水平

先合并A、B两列的因子水平,确保二者水平一致:

combined_levels <- union(levels(Data_Frame$A), levels(Data_Frame$B))
Data_Frame$A <- factor(Data_Frame$A, levels = combined_levels)
Data_Frame$B <- factor(Data_Frame$B, levels = combined_levels)

步骤2:转换value为数值型

将字符型的value转为数值,方便后续比较:

Data_Frame$value <- as.numeric(Data_Frame$value)

步骤3:生成配对唯一标识

对每行的A、B值按字母排序后拼接,让回文对生成相同的标识:

Data_Frame$pair_id <- apply(Data_Frame[, c("A", "B")], 1, function(x) paste(sort(x), collapse = "-"))

步骤4:分组筛选最小value行

按pair_id和effect分组,保留每组中value最小的行:

library(dplyr)
result <- Data_Frame %>%
  group_by(pair_id, effect) %>%
  filter(value == min(value)) %>%
  ungroup() %>%
  select(-pair_id) %>%
  distinct()

Python 实现方案

基于pandas库完成:

步骤1:构建数据框并转换value类型

import pandas as pd

df = pd.DataFrame({
    "A": ["a", "c", "b", "e", "g", "d", "f", "h"],
    "B": ["b", "d", "a", "f", "h", "c", "e", "g"],
    "value": ["0.3", "0.2", "0.1", "0.1", "0.5", "0.7", "0.8", "0.1"],
    "effect": ["123", "345", "123", "444", "123", "345", "444", "123"]
})

# 转换value为数值型
df["value"] = pd.to_numeric(df["value"])

步骤2:生成配对唯一标识

对每行A、B排序后拼接,生成回文对的统一标识:

df["pair_id"] = df.apply(lambda row: "-".join(sorted([row["A"], row["B"]])), axis=1)

步骤3:分组筛选最小value行

按pair_id和effect分组,保留每组中value最小的行:

result = df.groupby(["pair_id", "effect"]).apply(
    lambda x: x[x["value"] == x["value"].min()]
).reset_index(drop=True)

# 移除辅助列pair_id并去重
result = result.drop("pair_id", axis=1).drop_duplicates()

内容的提问来源于stack exchange,提问作者RJF

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 12:23:24