如何用PyDatatable实现类似R data.table中.SDcols的多列更新?
.SDcols in R data.table? I'm working with the Iris dataset and want to add "distance from mean" columns for all numeric columns. Currently I can do this by explicitly listing each column:
from datatable import fread, f, mean, update, abs iris_dt = fread("https://h2o-public-test-data.s3.amazonaws.com/smalldata/iris/iris.csv") iris_dt[:, update(C0_dist_from_mean = abs(f.C0 - mean(f.C0)), C1_dist_from_mean = abs(f.C1 - mean(f.C1)), C2_dist_from_mean = abs(f.C2 - mean(f.C2)), C3_dist_from_mean = abs(f.C3 - mean(f.C1)))]
But this approach hardcodes column names and isn't robust. In R's data.table, I can use .SDcols for flexible batch processing:
library(data.table) iris = fread("https://h2o-public-test-data.s3.amazonaws.com/smalldata/iris/iris.csv") cols = names(sapply(iris, class)[sapply(iris, class)=='numeric']) iris[, paste0(cols,"_dist_from_mean") := lapply(.SD, function(x) {abs(x-mean(x))}), .SDcols=cols]
I already know how to filter numeric columns in PyDatatable (like iris_dt[:, f[float]]), but I don't know how to implement batch update logic similar to R's .SDcols.
Solution in PyDatatable
Absolutely! You can replicate that flexible batch update behavior in PyDatatable by dynamically building your update expressions. Here's a step-by-step breakdown to achieve this:
Step 1: Grab your numeric column names
First, we need to identify which columns are numeric (float-type) in your datatable. This avoids hardcoding and automatically adapts if your dataset changes:
numeric_cols = [col for col in iris_dt.names if iris_dt[:, f[col]].stype.is_float]
Step 2: Build update expressions dynamically
Next, create a dictionary where each entry maps a new column name (original column + _dist_from_mean) to the expression that calculates the absolute distance from the column's mean:
update_exprs = { f"{col}_dist_from_mean": abs(f[col] - mean(f[col])) for col in numeric_cols }
(Note: Fixed a typo from your original code here—your C3 calculation was using the mean of C1, which this approach avoids entirely by referencing each column's own mean.)
Step 3: Run the batch update
Pass the unpacked dictionary to update()—PyDatatable will process all these expressions in a single efficient operation:
iris_dt[:, update(**update_exprs)]
Full Working Code
Here's the complete script you can run:
from datatable import fread, f, mean, update, abs # Load the Iris dataset iris_dt = fread("https://h2o-public-test-data.s3.amazonaws.com/smalldata/iris/iris.csv") # Identify numeric columns numeric_cols = [col for col in iris_dt.names if iris_dt[:, f[col]].stype.is_float] # Create update expressions for each numeric column update_exprs = { f"{col}_dist_from_mean": abs(f[col] - mean(f[col])) for col in numeric_cols } # Perform the batch update iris_dt[:, update(**update_exprs)] # Check the first few rows to verify print(iris_dt.head())
How this compares to R's .SDcols
This approach mirrors the flexibility of .SDcols:
- It automatically targets only numeric columns, no manual column listing needed
- It scales seamlessly if you add or remove numeric columns later
- The dictionary comprehension acts like the
lapply(.SD, ...)part, generating the required calculation for each column in the selected set
内容的提问来源于stack exchange,提问作者topchef

