You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

伪算法转Pandas代码求助:实现带重叠处理的数据集行选择函数

Solution: Pandas Implementation of Your Row Selection Logic

Let's build this function step by step, focusing on solving the overlap resolution that's stumping you. The core trick is using a set to track selected row indices (for fast duplicate checks) and building helper logic to fetch non-overlapping rows dynamically when needed.

First, Import Required Libraries

import pandas as pd
import numpy as np
import random

Helper Function: Fetch Next Unique Rows

This helper grabs rows from a column starting at a specific rank, skipping any already selected indices to avoid overlaps:

def get_next_unique_rows(df, col, start_rank, existing_indices, num_needed):
    # Sort the column in descending order and get all row indices
    sorted_indices = df[col].sort_values(ascending=False).index.tolist()
    # Filter out already selected indices, starting from the specified rank
    valid_candidates = [idx for idx in sorted_indices[start_rank:] if idx not in existing_indices]
    # Return exactly the number of rows needed (dataset has 100 rows, so we won't run out)
    return valid_candidates[:num_needed]

Main Function Implementation

This function handles all 5 cases of cols_to_use length, resolving both overlap1 and overlap2 events automatically:

def f(df, cols_to_use = ["X", "Y", "Z", "F", "V"]):
    # Define initial rows per column for each case
    rows_per_col = {5: 2, 4: 2, 3: 3, 2: 5, 1: 10}
    n_cols = len(cols_to_use)
    initial_rows = rows_per_col[n_cols]
    
    # Use a set to track selected indices (O(1) duplicate checks for speed)
    selected_indices = set()

    # Step 1: Handle initial row selection + resolve overlap1
    for col in cols_to_use:
        sorted_col_indices = df[col].sort_values(ascending=False).index.tolist()
        needed = initial_rows
        current_rank = 0
        
        # Keep fetching rows until we get the required number of unique ones
        while needed > 0 and current_rank < len(df):
            idx = sorted_col_indices[current_rank]
            if idx not in selected_indices:
                selected_indices.add(idx)
                needed -= 1
            current_rank += 1

    # Step 2: Handle supplementary rows + resolve overlap2
    if n_cols == 4:
        # Need 2 more rows: randomly pick 2 columns, fetch 1 new row each (starting at rank 2)
        random_cols = random.sample(cols_to_use, 2)
        for col in random_cols:
            new_idx = get_next_unique_rows(df, col, start_rank=2, existing_indices=selected_indices, num_needed=1)[0]
            selected_indices.add(new_idx)
    elif n_cols == 3:
        # Need 1 more row: randomly pick 1 column, fetch 1 new row (starting at rank 3)
        random_col = random.sample(cols_to_use, 1)[0]
        new_idx = get_next_unique_rows(df, random_col, start_rank=3, existing_indices=selected_indices, num_needed=1)[0]
        selected_indices.add(new_idx)

    # Sanity check to ensure we have exactly 10 unique rows
    assert len(selected_indices) == 10, f"Expected 10 unique rows, got {len(selected_indices)}"

    # Return the final filtered DataFrame (preserves original row order if needed)
    return df.loc[list(selected_indices), :]

How This Solves Your Overlap Problems

  • overlap1 Resolution: When selecting initial rows, if a row is already in the selected set, we automatically skip it and fetch the next highest-ranked row for that column until we get the required number of unique entries.
  • overlap2 Resolution: The helper function get_next_unique_rows ensures any supplementary rows don't overlap with already selected indices by checking against the set and fetching the next valid row if needed.
  • Efficiency: Using a set for index tracking makes duplicate checks nearly instant, which is much faster than repeatedly filtering a DataFrame.

Test with Your Sample Dataset

# Generate sample data
sample = pd.DataFrame(np.random.randint(0, 500, size=(100, 5)))
sample.columns = ["X", "Y", "Z", "F", "V"]

# Test the function
filtered_result = f(sample)
print(f"Filtered shape: {filtered_result.shape}")  # Output: (10, 5)
print(f"Unique rows: {filtered_result.index.nunique()}")  # Output: 10

Key Notes

  • The function uses descending order for "top" ranks (highest column values first). If you need ascending order, just change ascending=False to ascending=True in the sort calls.
  • The assertion is a sanity check—you can replace it with a custom exception or remove it if you prefer.
  • For the 1-column case, we directly take the top 10 rows since there's no possibility of overlap (each row is unique in the dataset).

内容的提问来源于stack exchange,提问作者Emil Mirzayev

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 20:44:08