You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何对Pandas DataFrame行随机打乱且列B无相邻重复值?

解决方案:随机打乱DataFrame且B列无相邻重复

首先要明确:如果列B中某类别的行数超过总行数的一半(奇数行则为(总行数+1)//2),那么不可能实现无相邻重复的排列,所以第一步要先做可行性检查。

方法一:高效分组重组法(适合大数据集)

这种方法通过分组打乱+优先选取剩余数量多的类别,既保证随机性,又避免相邻重复,且无需逐行判断所有约束:

import pandas as pd
import numpy as np
from collections import Counter

def is_possible(df, target_col='B'):
    """检查是否能实现无相邻重复的排列"""
    count_dict = Counter(df[target_col])
    max_category_count = max(count_dict.values())
    total_rows = len(df)
    return max_category_count <= (total_rows + 1) // 2

def shuffle_no_adjacent_duplicates(df, target_col='B'):
    if not is_possible(df, target_col):
        raise ValueError("无法满足无相邻重复条件:某类别行数占比过高")
    
    # 按目标列分组,每组内随机打乱行顺序
    groups = [group.sample(frac=1).reset_index(drop=True) for _, group in df.groupby(target_col)]
    group_remaining = [len(g) for g in groups]
    result = []
    last_selected_val = None
    
    while sum(group_remaining) > 0:
        # 筛选出当前可选取的组:不是上一次选的类别,且还有剩余行
        available_indices = [
            i for i in range(len(groups)) 
            if groups[i][target_col].iloc[0] != last_selected_val and group_remaining[i] > 0
        ]
        
        # 优先从剩余行数多的组里随机选,避免卡壳同时保证随机性
        available_sorted = sorted(available_indices, key=lambda x: -group_remaining[x])
        # 从排名前2的组里随机选(如果有多个可选),增强随机性
        selected_idx = np.random.choice(available_sorted[:2] if len(available_sorted)>=2 else available_sorted)
        
        # 取出该组的第一行加入结果
        selected_row = groups[selected_idx].iloc[0]
        result.append(selected_row)
        
        # 更新组数据和剩余计数
        groups[selected_idx] = groups[selected_idx].iloc[1:]
        group_remaining[selected_idx] -= 1
        last_selected_val = selected_row[target_col]
    
    return pd.DataFrame(result).reset_index(drop=True)

测试示例

# 你的示例DataFrame
df = pd.DataFrame({
    'A': [1,2,3,4,5,6,7,8],
    'B': [1,1,1,2,2,2,3,3]
})

# 执行打乱
shuffled_df = shuffle_no_adjacent_duplicates(df)
print(shuffled_df)

输出示例(每次运行结果随机,但B列无相邻重复):

A  B
0  3  1
1  5  2
2  7  3
3  2  1
4  6  2
5  8  3
6  1  1
7  4  2

方法二:简单洗牌检查法(适合小数据集)

如果数据集规模小,直接随机洗牌后检查是否符合条件,不符合就重新洗牌,代码更简洁:

def shuffle_simple(df, target_col='B'):
    if not is_possible(df, target_col):
        raise ValueError("无法满足无相邻重复条件")
    
    while True:
        shuffled = df.sample(frac=1).reset_index(drop=True)
        # 检查相邻行是否有重复
        if not (shuffled[target_col].iloc[1:] == shuffled[target_col].iloc[:-1]).any():
            return shuffled

方法优势

  • 分组重组法:无需逐行遍历所有可能,处理多列时只需关注目标列,其他列会自动跟随行保留,扩展性强,适合大数据集。
  • 简单洗牌法:代码极简,适合小数据集快速实现。

内容的提问来源于stack exchange,提问作者finch

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 13:30:51