如何提取缺失值占比低于20%的数据集列名并生成可用列表?
Hey there! Great job writing that missing value percentage function—let's build on that to get the column names you need, no manual copying required.
Step 1: Filter the missing rate vector to get valid column names
Your x variable is a named vector where each name corresponds to a column from dt2, and the value is its missing percentage. To grab only the columns with less than 20% missing values, just subset x to keep entries where the value is <20, then pull their names:
# Get column names where missing rate < 20% keep_cols <- names(x[x < 20])
Step 2: Create your new dataset directly
Now you can use this keep_cols vector to slice your original dataset into the reduced version you want. Depending on whether you're using a base data frame or data.table (since your object is named dt2, I suspect this might be your case), here are the two common approaches:
Base R data frame:
new_dt2 <- dt2[, keep_cols, drop = FALSE]The
drop = FALSEensures that even if only one column is kept, it stays as a data frame instead of converting to a vector.data.table:
new_dt2 <- dt2[, ..keep_cols]The
..prefix tellsdata.tableto look for the variable in the global environment instead of treating it as a column name.
Quick check to verify
If you want to double-check that you got the right columns, you can print the missing rates for the kept columns:
x[keep_cols]
That's it! No more manual copying—this workflow will give you a reusable list of columns and a clean new dataset in one go.
内容的提问来源于stack exchange,提问作者Aaron England

