You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于优先级的Pandas DataFrame抽样:寻求可扩展的优化实现方案

Prioritized Sampling in Pandas: Scalable Approach

Great question! This kind of tiered prioritized sampling comes up often, and you’re right—there’s a far cleaner, scalable way to handle it than writing a mess of conditional checks for each priority level. Here’s how to implement it in a way that works no matter how many priority groups you have:

Core Idea

We’ll process groups in priority order (A → B → C, or whatever your hierarchy is), taking as many samples as possible from each group first—either the entire group if it’s smaller than the remaining number of samples we need, or just the remaining number we need. Once we hit our target sample count, we stop.

Step-by-Step Code Implementation

Let’s walk through it with your sample data (I added a random seed for reproducibility):

import random
import pandas as pd

# Generate test data (same as your code, plus seed for consistency)
random.seed(42)
col1 = random.choices([1,2,3,4,5], k=50)
col2 = random.choices(['A','B','C'], k=50)
df = pd.DataFrame({'values': col1, 'priority': col2})

# Define your priority hierarchy (easily extendable—add 'D', 'E', etc., here)
priority_order = ['A', 'B', 'C']
target_sample_size = 25

# Split the DataFrame into groups matching our priority order
priority_groups = [df[df['priority'] == p] for p in priority_order]

# Collect our sampled data
selected_samples = []
remaining_samples = target_sample_size

for group in priority_groups:
    if remaining_samples <= 0:
        break  # We've hit our target, no need to process further groups
    
    # Calculate how many samples to take from this group
    num_to_take = min(len(group), remaining_samples)
    
    # Randomly select the required number of samples
    sampled_group = group.sample(n=num_to_take, random_state=42)
    selected_samples.append(sampled_group)
    
    # Update how many samples we still need
    remaining_samples -= num_to_take

# Combine all selected samples into a single DataFrame
final_sample = pd.concat(selected_samples).reset_index(drop=True)

# Verify the result
print(f"Total sampled rows: {len(final_sample)}")
print("\nPriority distribution in sample:")
print(final_sample['priority'].value_counts())

Why This Works (and Why It’s Better)

  • Scalable: If you later add more priority levels (like 'D' or 'E'), you just add them to the priority_order list—no need to rewrite conditional logic.
  • Clean & Readable: No nested if-elif chains. The logic flows linearly, making it easy to debug or modify.
  • Handles All Edge Cases: It automatically covers every scenario you mentioned:
    • If 'A' has 25+ samples: takes exactly 25 from 'A'
    • If 'A' has 20, 'B' has 30: takes all 20 from 'A', then 5 from 'B'
    • If 'A' has 10, 'B' has 10: takes all 10 from 'A' and 10 from 'B', then 5 from 'C'

Optional: Using groupby for Cleaner Grouping

If you prefer using groupby instead of list comprehensions, you can adjust the grouping step like this (just make sure to preserve the priority order):

# Using groupby (sort=False ensures we keep the original priority order)
priority_groups = []
for p in priority_order:
    group = df.groupby('priority', sort=False).get_group(p)
    priority_groups.append(group)

This does the same thing as the list comprehension—it’s just a matter of personal preference.

内容的提问来源于stack exchange,提问作者Phil

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:21:22