You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提取缺失值占比低于20%的数据集列名并生成可用列表?

Extract Columns with Missing Rate < 20%

Hey there! Great job writing that missing value percentage function—let's build on that to get the column names you need, no manual copying required.

Step 1: Filter the missing rate vector to get valid column names

Your x variable is a named vector where each name corresponds to a column from dt2, and the value is its missing percentage. To grab only the columns with less than 20% missing values, just subset x to keep entries where the value is <20, then pull their names:

# Get column names where missing rate < 20%
keep_cols <- names(x[x < 20])

Step 2: Create your new dataset directly

Now you can use this keep_cols vector to slice your original dataset into the reduced version you want. Depending on whether you're using a base data frame or data.table (since your object is named dt2, I suspect this might be your case), here are the two common approaches:

  • Base R data frame:

    new_dt2 <- dt2[, keep_cols, drop = FALSE]
    

    The drop = FALSE ensures that even if only one column is kept, it stays as a data frame instead of converting to a vector.

  • data.table:

    new_dt2 <- dt2[, ..keep_cols]
    

    The .. prefix tells data.table to look for the variable in the global environment instead of treating it as a column name.

Quick check to verify

If you want to double-check that you got the right columns, you can print the missing rates for the kept columns:

x[keep_cols]

That's it! No more manual copying—this workflow will give you a reusable list of columns and a clean new dataset in one go.

内容的提问来源于stack exchange,提问作者Aaron England

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 08:01:09