You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用PyDatatable实现类似R data.table中.SDcols的多列更新?

How to Batch Update Columns in PyDatatable Like .SDcols in R data.table?

I'm working with the Iris dataset and want to add "distance from mean" columns for all numeric columns. Currently I can do this by explicitly listing each column:

from datatable import fread, f, mean, update, abs
iris_dt = fread("https://h2o-public-test-data.s3.amazonaws.com/smalldata/iris/iris.csv")
iris_dt[:, update(C0_dist_from_mean = abs(f.C0 - mean(f.C0)), 
                  C1_dist_from_mean = abs(f.C1 - mean(f.C1)), 
                  C2_dist_from_mean = abs(f.C2 - mean(f.C2)), 
                  C3_dist_from_mean = abs(f.C3 - mean(f.C1)))]

But this approach hardcodes column names and isn't robust. In R's data.table, I can use .SDcols for flexible batch processing:

library(data.table)
iris = fread("https://h2o-public-test-data.s3.amazonaws.com/smalldata/iris/iris.csv")
cols = names(sapply(iris, class)[sapply(iris, class)=='numeric'])
iris[, paste0(cols,"_dist_from_mean") := lapply(.SD, function(x) {abs(x-mean(x))}), .SDcols=cols]

I already know how to filter numeric columns in PyDatatable (like iris_dt[:, f[float]]), but I don't know how to implement batch update logic similar to R's .SDcols.


Solution in PyDatatable

Absolutely! You can replicate that flexible batch update behavior in PyDatatable by dynamically building your update expressions. Here's a step-by-step breakdown to achieve this:

Step 1: Grab your numeric column names

First, we need to identify which columns are numeric (float-type) in your datatable. This avoids hardcoding and automatically adapts if your dataset changes:

numeric_cols = [col for col in iris_dt.names if iris_dt[:, f[col]].stype.is_float]

Step 2: Build update expressions dynamically

Next, create a dictionary where each entry maps a new column name (original column + _dist_from_mean) to the expression that calculates the absolute distance from the column's mean:

update_exprs = {
    f"{col}_dist_from_mean": abs(f[col] - mean(f[col]))
    for col in numeric_cols
}

(Note: Fixed a typo from your original code here—your C3 calculation was using the mean of C1, which this approach avoids entirely by referencing each column's own mean.)

Step 3: Run the batch update

Pass the unpacked dictionary to update()—PyDatatable will process all these expressions in a single efficient operation:

iris_dt[:, update(**update_exprs)]

Full Working Code

Here's the complete script you can run:

from datatable import fread, f, mean, update, abs

# Load the Iris dataset
iris_dt = fread("https://h2o-public-test-data.s3.amazonaws.com/smalldata/iris/iris.csv")

# Identify numeric columns
numeric_cols = [col for col in iris_dt.names if iris_dt[:, f[col]].stype.is_float]

# Create update expressions for each numeric column
update_exprs = {
    f"{col}_dist_from_mean": abs(f[col] - mean(f[col]))
    for col in numeric_cols
}

# Perform the batch update
iris_dt[:, update(**update_exprs)]

# Check the first few rows to verify
print(iris_dt.head())

How this compares to R's .SDcols

This approach mirrors the flexibility of .SDcols:

  • It automatically targets only numeric columns, no manual column listing needed
  • It scales seamlessly if you add or remove numeric columns later
  • The dictionary comprehension acts like the lapply(.SD, ...) part, generating the required calculation for each column in the selected set

内容的提问来源于stack exchange,提问作者topchef

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 15:57:32