You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按指定列组合去重并保留rykkedage最小的行?

Hey Louise, great question! Let's break down how to handle both scenarios you're asking about—first keeping only the row with the smallest rykkedage for duplicate column combinations, then extending that to keep all rows where rykkedage matches the group minimum (even if there are ties). We'll use Python's pandas library since it's the go-to tool for this kind of data cleaning.

1. Keep Rows with the Smallest rykkedage (Single Row per Group)

If you only want one row per duplicate combination (the one with the smallest rykkedage), here's a concise way to do it:

import pandas as pd

# Load your data (adjust the file path/format as needed)
df = pd.read_csv("your_dataset.csv")

# Define the columns that determine a duplicate row
duplicate_key = ["ID", "BilagNr", "Henstand", "Aftale", "Belob", "RP", "Pos", "Dps", "Udlign"]

# Group by the duplicate key and keep the row with the smallest rykkedage
filtered_df = df.loc[df.groupby(duplicate_key)["rykkedage"].idxmin()]

# Reset the index for cleaner output (optional)
filtered_df = filtered_df.reset_index(drop=True)

How this works:

  • groupby(duplicate_key) clusters rows that are identical across your specified columns.
  • ["rykkedage"].idxmin() finds the index of the row in each group with the smallest rykkedage value.
  • loc[...] extracts those specific rows from the original DataFrame.

2. Keep All Rows with the Smallest rykkedage (Including Ties)

If you want to retain all rows where rykkedage equals the group minimum (even if multiple rows share that smallest value), adjust the code like this:

import pandas as pd

df = pd.read_csv("your_dataset.csv")
duplicate_key = ["ID", "BilagNr", "Henstand", "Aftale", "Belob", "RP", "Pos", "Dps", "Udlign"]

# Add a temporary column with the minimum rykkedage for each group
df["group_min_rykkedage"] = df.groupby(duplicate_key)["rykkedage"].transform("min")

# Filter rows where rykkedage matches the group's minimum
filtered_with_ties_df = df[df["rykkedage"] == df["group_min_rykkedage"]]

# Remove the temporary column (optional but clean)
filtered_with_ties_df = filtered_with_ties_df.drop(columns=["group_min_rykkedage"])
filtered_with_ties_df = filtered_with_ties_df.reset_index(drop=True)

How this works:

  • transform("min") calculates the smallest rykkedage for each group and assigns that value to every row in the group (so we can compare individual rows to the group minimum).
  • The final filter keeps only rows where rykkedage matches the group's minimum value, preserving all ties.

Example Data & Results

Let's use a sample dataset to show how both solutions work:

Original Data

IDBilagNrHenstandAftaleBelobRPPosDpsUdlignrykkedage
11001H1A1500R1P1D1U12
11001H1A1500R1P1D1U11
11001H1A1500R1P1D1U11
21002H2A2300R2P2D2U23
21002H2A2300R2P2D2U22

Result: Single Row per Group (Minimum rykkedage)

IDBilagNrHenstandAftaleBelobRPPosDpsUdlignrykkedage
11001H1A1500R1P1D1U11
21002H2A2300R2P2D2U22

Result: All Tied Minimum Rows

IDBilagNrHenstandAftaleBelobRPPosDpsUdlignrykkedage
11001H1A1500R1P1D1U11
11001H1A1500R1P1D1U11
21002H2A2300R2P2D2U22

内容的提问来源于stack exchange,提问作者Louise Sørensen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 06:47:47