You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R中100行×10^7列data.table高效行筛选方案咨询

Efficient Row Filtering for a 100x10^7 data.table

Hey, let's break this down—your bottleneck isn't the row filtering logic itself, it's the extreme number of columns (10^7). Even with just 100 rows, copying subsets of that many columns creates massive memory and I/O overhead, which is why your current code is taking 20+ minutes. Below are the best tailored solutions, ordered by priority:

Since your data has far fewer rows than columns, transposing it turns the problem into a trivial column-filtering task (instead of row-filtering a wide table), then you can transpose back to get your desired train/test splits. Here's how to implement it properly:

First, separate your status column from the feature columns to avoid mixing it up during transposition:

# Split status and feature columns
status_df <- df[, .(status)]
feature_df <- df[, !"status"]

# Transpose the feature data (now features are rows, samples are columns)
feature_transposed <- transpose(feature_df, keep.names = "feature_id")

# Generate your train index as before (using the original status vector)
trainIndex <- caret::createDataPartition(df$status, p = .9, list = FALSE, times = 1)

# Filter columns in the transposed data (corresponds to filtering rows in original)
feature_train_transposed <- feature_transposed[, ..trainIndex]
feature_test_transposed <- feature_transposed[, -..trainIndex]

# Transpose back and reattach the status column
outerTrain <- cbind(status_df[trainIndex], transpose(feature_train_transposed, make.names = "feature_id"))
outerTest <- cbind(status_df[-trainIndex], transpose(feature_test_transposed, make.names = "feature_id"))

This works because transposing reduces the number of columns from 10^7 to 100. Filtering 100 columns is negligible compared to copying 10^7 vectors for row subsets.

2. Use Matrix Storage (If Data Types Are Uniform)

If all your feature columns are the same data type (e.g., numeric), converting the feature set to a matrix will drastically speed up row filtering. Matrices store data in contiguous memory blocks, so subsetting rows is a single block copy instead of copying 10^7 separate vectors like data.table does:

# Extract status and convert features to matrix
status_vec <- df$status
feature_matrix <- as.matrix(df[, !"status"])

# Filter rows directly on the matrix
train_matrix <- feature_matrix[trainIndex, ]
test_matrix <- feature_matrix[-trainIndex, ]

# Convert back to data.table if needed
outerTrain <- data.table(status = status_vec[trainIndex], train_matrix)
outerTest <- data.table(status = status_vec[-trainIndex], test_matrix)

Note: This only works if all features share the same data type—mixing types will force coercion, which may lose information or introduce errors.

3. Minimize Copy Overhead (Last Resort)

If transposing or using a matrix isn't feasible, you can try to reduce unnecessary copies, but this will still be slower than the above methods. One trick is to use logical indexing instead of integer indices, though the gains are minimal here:

# Convert trainIndex to a logical vector
train_logical <- logical(nrow(df))
train_logical[trainIndex] <- TRUE

# Subset using the logical vector
outerTrain <- df[train_logical]
outerTest <- df[!train_logical]

Alternatively, if you don't need to keep the original df, you can first extract the training set and then modify the original table in place to get the test set (though this is risky if you need the original data later):

outerTrain <- df[trainIndex]
df <- df[-trainIndex] # Overwrites original df with test data

Why Your Original Code Is So Slow

To clarify the root cause: data.table stores each column as an independent vector. When you run df[trainIndex], R has to create 10^7 new vectors (each with 90 elements for the training set) and assemble them into a new data.table. That's 9e8 individual element copies—an enormous amount of work. Transposing or using a matrix cuts this work down by orders of magnitude.


内容的提问来源于stack exchange,提问作者Keshav M

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 03:44:13