Julia中如何无需for循环将向量的向量转换为DataFrame
Julia嵌套向量转指定尺寸DataFrame实现方案
问题描述
- 已通过如下代码生成名为
all_arrays的嵌套向量,其长度为1000,每个内部元素为长度17的Float64类型向量:
using DataFrames using StatsBase list_of_numbers = 1:17 all_arrays = [zeros(Float64, (17,)) for i in 1:1000] round = 1 while round != 1001 random_array = StatsBase.sample(1:17 , length(list_of_numbers)) random_array = random_array/sum(random_array) if (0.0 in random_array) || (random_array in all_arrays) continue end all_arrays[round] = random_array round += 1 println(round) end
- 执行
size(all_arrays)返回结果为(1000,),目标是将其转换为1000行、17列的DataFrame结构。 - 当前采用的实现方案需要先预分配全零初始化的DataFrame,再通过显式for循环逐行赋值:
df = DataFrames.DataFrame(zeros(1000,17) , :auto) for idx in 1:length(all_arrays) df[idx , :] = all_arrays[idx] end
- 预期目标:找到更简洁的实现方式,不需要手动编写for循环,也不需要预先构建初始化的空/全零DataFrame。
最简实现
不需要预分配对象、不需要手动写循环,一行代码即可完成转换:
df = DataFrame(permutedims(reduce(hcat, all_arrays)), :auto)
实现逻辑说明
reduce(hcat, all_arrays):将所有长度为17的子向量按列横向拼接,得到尺寸为17行×1000列的连续内存矩阵permutedims:对拼接后的矩阵做转置操作,得到1000行×17列的目标尺寸矩阵- 直接将转置后的矩阵传入
DataFrame构造函数,参数:auto会自动生成从x1到x17的默认列名
该实现性能显著优于逐行循环赋值的方案,因为全程是连续内存的批量操作,避免了循环逐行修改DataFrame结构产生的额外开销。
如果需要自定义列名,可以直接在构造时传入列名列表,例如将列命名为feat_1到feat_17:
col_names = Symbol.( "feat_" .* string.(1:17) ) df = DataFrame(permutedims(reduce(hcat, all_arrays)), col_names)
内容的提问来源于stack exchange,提问作者Shayan
相关产品推荐
相关产品推荐

