You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何对超大规模PyArrow数据集进行有放回随机抽样

大规模Arrow数据集有放回抽样及后续随机森林处理方案

一、PyArrow+Hugging Face Dataset 有放回抽样实现

直接转Pandas会因内存不足崩溃,全程用内存友好的PyArrow或Hugging Face Dataset方案:

方法1:基于PyArrow Table的可靠有放回抽样

先将Hugging Face Dataset转为PyArrow Table(无需全量加载到内存),再生成允许重复的随机索引完成抽样:

from datasets import Dataset
import numpy as np

# 加载Arrow数据集(延迟加载,不占满内存)
dataset = Dataset.from_file("embeddings_job/combined_embeddings_small/data-00000-of-00001.arrow")
pyarrow_table = dataset.to_table()

# 执行20次有放回抽样,每次100行
for sample_round in range(20):
    # np.random.randint允许生成重复索引,实现有放回
    random_indices = np.random.randint(0, pyarrow_table.num_rows, size=100)
    # 抽样,支持重复索引,保留原数据的行信息
    sampled_table = pyarrow_table.select(random_indices)
    # 转为小批量Pandas DataFrame(100行无内存压力)
    sampled_df = sampled_table.to_pandas()
    # 后续处理(如喂随机森林)
    process_sample(sampled_df)

方法2:Hugging Face Dataset原生抽样(需版本支持)

若你的Dataset版本支持replace参数,可直接用shuffle实现有放回抽样:

from datasets import Dataset
import numpy as np

dataset = Dataset.from_file("embeddings_job/combined_embeddings_small/data-00000-of-00001.arrow")

for sample_round in range(20):
    # replace=True开启有放回抽样,每次取100行
    sampled_dataset = dataset.shuffle(seed=np.random.randint(0, 10000), replace=True).select(range(100))
    sampled_df = sampled_dataset.to_pandas()
    process_sample(sampled_df)

注:若shuffle不支持replace参数,优先用方法1的随机索引方案。

二、终端命令的修正与说明

你提供的终端命令可以实现有放回抽样,但需调整参数匹配需求:

# 20次抽样,每次100行,-r开启有放回,输出到sampled_1.txt至sampled_20.txt
for i in {1..20}; do shuf -n 100 -r embeddings_job/combined_embeddings_small/data-00000-of-00001.arrow > sampled_$i.txt; done

注意:shuf是按文本行处理文件的,若你的Arrow是二进制格式,此方法会失效,建议优先用Python代码方案。

三、抽样数据喂随机森林的最佳实践

每次抽样仅100行,完全可以转为Pandas DataFrame后用Scikit-learn的随机森林模型:

from sklearn.ensemble import RandomForestClassifier

def process_sample(sampled_df):
    # 分离特征与标签(假设标签列名为'label')
    X = sampled_df.drop('label', axis=1)
    y = sampled_df['label']
    
    # 初始化并训练随机森林(可根据需求调整参数)
    rf_model = RandomForestClassifier(n_estimators=100, random_state=42)
    rf_model.fit(X, y)
    
    # 后续评估、预测或模型保存
    # ...

如果需要融合多次抽样的训练结果,可收集每次训练的模型,最终通过投票、平均预测等方式集成。

四、确保子集索引不重置的要点

  • 抽样时不要调用reset_index(),PyArrow Table和Hugging Face Dataset的抽样结果会保留原数据的行索引。
  • 若原数据集无显式索引列,加载时可主动添加:
# 给数据集添加原索引列,抽样后可追溯原行位置
dataset = dataset.add_column("original_index", list(range(len(dataset))))

内容的提问来源于stack exchange,提问作者youtube

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 19:23:13