You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

EDA:含65%以上NA值的列能否删除?原因及替代处理方式

Handling Columns with >65% Missing Values in EDA

Great question—this is one of those call-it-as-you-see-it decisions we face all the time during exploratory data analysis, and there’s no hard rule, but let’s walk through when deleting makes sense, when it doesn’t, and what to do instead.

When You Can Safely Delete the Column

  • If the column has no meaningful business context and can’t contribute to your analysis or model:
    Think of a random "optional notes" field where 90% of entries are blank, and the remaining 10% are gibberish or irrelevant. Keeping it just adds noise, slows down computations, and gives you no actionable insights.
  • If the missing values are random and carry no signal:
    Like a system glitch that caused an entire column to fail to collect data for months. There’s no story behind the gaps—they’re just bad data. Ditching the column is the cleanest move.

When You Shouldn’t Delete the Column (and What to Do Instead)

This is where the interesting stuff happens—high missing rates often hide useful information:

  • The missing values themselves are a key signal:
    Suppose you’re analyzing e-commerce users, and the "business tax ID" column has 70% missing values. Those gaps aren’t random—they mean most users are individual shoppers, not businesses. Deleting this column would erase a critical way to segment your user base.
    Fix: Convert the column into a binary feature: 1 if the value exists, 0 if it’s missing. Now you’re leveraging the missingness as data.
  • The column is a core business field with high predictive power (even with gaps):
    In credit scoring, an "asset verification" column might have 65% missing values, but the users who do provide this info are far less likely to default. Ditching it would throw away a strong predictor.
    Fixes:
    • Targeted imputation: Use the median/mode of non-missing values, or build a small predictive model using other related columns (like income, credit score) to fill in gaps.
    • Use models that handle missing natively: Tree-based models like XGBoost or Random Forest don’t require pre-imputation—they learn how to treat missing values as a separate category automatically.
  • You expect to get more data later:
    If you know a batch of historical data is coming next week that will fill in most of the gaps, hold off on deleting. No need to throw away a potentially useful column prematurely.

内容的提问来源于stack exchange,提问作者Mirza Hashim

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 03:56:49