伪算法转Pandas代码求助:实现带重叠处理的数据集行选择函数
Solution: Pandas Implementation of Your Row Selection Logic
Let's build this function step by step, focusing on solving the overlap resolution that's stumping you. The core trick is using a set to track selected row indices (for fast duplicate checks) and building helper logic to fetch non-overlapping rows dynamically when needed.
First, Import Required Libraries
import pandas as pd import numpy as np import random
Helper Function: Fetch Next Unique Rows
This helper grabs rows from a column starting at a specific rank, skipping any already selected indices to avoid overlaps:
def get_next_unique_rows(df, col, start_rank, existing_indices, num_needed): # Sort the column in descending order and get all row indices sorted_indices = df[col].sort_values(ascending=False).index.tolist() # Filter out already selected indices, starting from the specified rank valid_candidates = [idx for idx in sorted_indices[start_rank:] if idx not in existing_indices] # Return exactly the number of rows needed (dataset has 100 rows, so we won't run out) return valid_candidates[:num_needed]
Main Function Implementation
This function handles all 5 cases of cols_to_use length, resolving both overlap1 and overlap2 events automatically:
def f(df, cols_to_use = ["X", "Y", "Z", "F", "V"]): # Define initial rows per column for each case rows_per_col = {5: 2, 4: 2, 3: 3, 2: 5, 1: 10} n_cols = len(cols_to_use) initial_rows = rows_per_col[n_cols] # Use a set to track selected indices (O(1) duplicate checks for speed) selected_indices = set() # Step 1: Handle initial row selection + resolve overlap1 for col in cols_to_use: sorted_col_indices = df[col].sort_values(ascending=False).index.tolist() needed = initial_rows current_rank = 0 # Keep fetching rows until we get the required number of unique ones while needed > 0 and current_rank < len(df): idx = sorted_col_indices[current_rank] if idx not in selected_indices: selected_indices.add(idx) needed -= 1 current_rank += 1 # Step 2: Handle supplementary rows + resolve overlap2 if n_cols == 4: # Need 2 more rows: randomly pick 2 columns, fetch 1 new row each (starting at rank 2) random_cols = random.sample(cols_to_use, 2) for col in random_cols: new_idx = get_next_unique_rows(df, col, start_rank=2, existing_indices=selected_indices, num_needed=1)[0] selected_indices.add(new_idx) elif n_cols == 3: # Need 1 more row: randomly pick 1 column, fetch 1 new row (starting at rank 3) random_col = random.sample(cols_to_use, 1)[0] new_idx = get_next_unique_rows(df, random_col, start_rank=3, existing_indices=selected_indices, num_needed=1)[0] selected_indices.add(new_idx) # Sanity check to ensure we have exactly 10 unique rows assert len(selected_indices) == 10, f"Expected 10 unique rows, got {len(selected_indices)}" # Return the final filtered DataFrame (preserves original row order if needed) return df.loc[list(selected_indices), :]
How This Solves Your Overlap Problems
- overlap1 Resolution: When selecting initial rows, if a row is already in the selected set, we automatically skip it and fetch the next highest-ranked row for that column until we get the required number of unique entries.
- overlap2 Resolution: The helper function
get_next_unique_rowsensures any supplementary rows don't overlap with already selected indices by checking against the set and fetching the next valid row if needed. - Efficiency: Using a set for index tracking makes duplicate checks nearly instant, which is much faster than repeatedly filtering a DataFrame.
Test with Your Sample Dataset
# Generate sample data sample = pd.DataFrame(np.random.randint(0, 500, size=(100, 5))) sample.columns = ["X", "Y", "Z", "F", "V"] # Test the function filtered_result = f(sample) print(f"Filtered shape: {filtered_result.shape}") # Output: (10, 5) print(f"Unique rows: {filtered_result.index.nunique()}") # Output: 10
Key Notes
- The function uses descending order for "top" ranks (highest column values first). If you need ascending order, just change
ascending=Falsetoascending=Truein the sort calls. - The assertion is a sanity check—you can replace it with a custom exception or remove it if you prefer.
- For the 1-column case, we directly take the top 10 rows since there's no possibility of overlap (each row is unique in the dataset).
内容的提问来源于stack exchange,提问作者Emil Mirzayev
相关产品推荐
相关产品推荐

