如何用指定随机种子可复现地打乱DataFrame的行和列?
问题
希望使用指定随机种子打乱DataFrame的行和列,但设置Random.seed!(123)后,每次运行sample.获取打乱索引的结果都不相同,代码示例如下:
using Random, DataFrames, StatsBase Random.seed!(123) df = DataFrame( col1 = [1, 2, 3], col2 = [4, 5, 6] ); idx_row, idx_col = sample.( [1:size(df, 1), 1:size(df, 2)], [length(1:size(df, 1)), length(1:size(df, 2))], replace=false ) # 第一次输出:[1, 2, 3] 和 [2, 1] idx_row, idx_col = sample.( [1:size(df, 1), 1:size(df, 2)], [length(1:size(df, 1)), length(1:size(df, 2))], replace=false ) # 第二次输出:[2, 1, 3] 和 [2, 1]
需要实现可复现的DataFrame行和列打乱。
解决方法
原因分析
问题出在每次调用随机函数都会消耗全局随机数流的状态,第一次sample.执行后,全局随机数生成器的状态已经改变,第二次调用自然会得到不同结果。要实现可复现性,需要确保每次打乱操作都从相同的随机种子状态开始,或者显式控制随机数生成器。
方案1:每次打乱前重新设置种子
如果需要多次重复相同的打乱结果,每次执行打乱操作前重新设置随机种子即可:
using Random, DataFrames, StatsBase df = DataFrame( col1 = [1, 2, 3], col2 = [4, 5, 6] ); # 第一次打乱 Random.seed!(123) idx_row = sample(1:size(df, 1), size(df, 1), replace=false) idx_col = sample(1:size(df, 2), size(df, 2), replace=false) shuffled_df = df[idx_row, idx_col] # 结果固定:行顺序[1,2,3],列顺序[2,1] # 第二次打乱(重新设种子,结果和第一次一致) Random.seed!(123) idx_row = sample(1:size(df, 1), size(df, 1), replace=false) idx_col = sample(1:size(df, 2), size(df, 2), replace=false) shuffled_df2 = df[idx_row, idx_col] # shuffled_df 和 shuffled_df2 完全相同
方案2:使用独立的随机数生成器(RNG)
避免影响全局随机数流,可以创建独立的RNG实例,每次使用同一个RNG来保证可复现性:
using Random, DataFrames, StatsBase df = DataFrame( col1 = [1, 2, 3], col2 = [4, 5, 6] ); # 创建带固定种子的独立RNG rng = MersenneTwister(123) idx_row = sample(rng, 1:size(df, 1), size(df, 1), replace=false) idx_col = sample(rng, 1:size(df, 2), size(df, 2), replace=false) shuffled_df = df[idx_row, idx_col] # 若要重复相同结果,重新创建相同种子的RNG即可 rng2 = MersenneTwister(123) idx_row2 = sample(rng2, 1:size(df, 1), size(df, 1), replace=false) idx_col2 = sample(rng2, 1:size(df, 2), size(df, 2), replace=false) shuffled_df2 = df[idx_row2, idx_col2]
方案3:直接使用shuffle函数(更简洁)
因为你是对整个行/列索引做无替换全采样,本质就是打乱顺序,用shuffle更直接,同样可以通过种子或独立RNG控制:
using Random, DataFrames df = DataFrame( col1 = [1, 2, 3], col2 = [4, 5, 6] ); # 全局种子方式 Random.seed!(123) shuffled_df = df[shuffle(1:nrow(df)), shuffle(1:ncol(df))] # 独立RNG方式 rng = MersenneTwister(123) shuffled_df = df[shuffle(rng, 1:nrow(df)), shuffle(rng, 1:ncol(df))]
内容的提问来源于stack exchange,提问作者Shayan
相关产品推荐
相关产品推荐

