You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas含缺失值手机号列转np.int64并保留缺失值需求问询

Solution for Converting Phone Number Column to Integer with Preserved Missing Values

Got it, let's fix this properly! The core problem here is that standard numpy integer types (like np.int64) can't handle missing values (NaN), which is why your column defaulted to float in the first place. To keep integer formatting and let count() correctly ignore missing values, you need to use pandas' nullable integer dtype (Int64 with a capital I).

Step-by-Step Solution

  1. Convert the column to pandas' nullable integer type. This will preserve NaN for missing values while converting valid phone numbers to integers:
import pandas as pd

# Apply the conversion to your column
rtl_one['HOME_PHONE_1'] = rtl_one['HOME_PHONE_1'].astype('Int64')

If your column has any unexpected non-numeric values (though your sample data only has floats/NaN), you can first clean it with pd.to_numeric to coerce invalid entries to NaN:

rtl_one['HOME_PHONE_1'] = pd.to_numeric(rtl_one['HOME_PHONE_1'], errors='coerce').astype('Int64')

Verify the Results

  • Check the column dtype to confirm it's now a nullable integer:
    print(rtl_one['HOME_PHONE_1'].dtype)
    # Output: Int64
    
  • Run count() to confirm missing values are ignored:
    print(rtl_one.count())
    # Output:
    # DW_ATH_CHAIN_NBR    7
    # DW_ATH_SEQ_NBR      7
    # HOME_PHONE_1        4
    # dtype: int64
    
  • View the DataFrame, and you'll see phone numbers as clean integers with NaN for missing entries:
    DW_ATH_CHAIN_NBR DW_ATH_SEQ_NBR HOME_PHONE_1
    0               11              0    5026372971
    1               12              0    7348561782
    2              105              0    9038961433
    3              123              0    9205853614
    4              101              1          NaN
    5               10              1          NaN
    6             1018              1          NaN
    

Why This Works

Pandas' Int64 dtype is built specifically for integer columns with missing values. Unlike np.int64, it supports NaN without forcing the column to float, and count() will automatically skip NaN values just like it does with float columns. This avoids the two issues you hit earlier: filling with 0 (which skews your count stats) or converting to string (which leaves unwanted .0 suffixes).

内容的提问来源于stack exchange,提问作者DPs

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 03:43:36