不使用Pandas/Numpy,如何转换DataFrame列数据类型并替换空值?
解决CSV列类型转换与空值替换问题
直接用pandas处理整列数据效率最高,不用手动遍历列表。按你的需求,步骤如下:
1. 导入pandas并读取CSV
import pandas as pd # 读取CSV时,把空字符串识别为NaN df = pd.read_csv('dataset.csv', na_values=[''])
2. 按需求转换列类型
先明确各列的目标类型:
- 保持字符串类型:
last、first、sex - 转为浮点型:
age、fare - 转为整数型:
sibsp、pclass、survived
执行转换:
# 转换浮点型,无法转换的值设为NaN df['age'] = pd.to_numeric(df['age'], errors='coerce').astype('float') df['fare'] = pd.to_numeric(df['fare'], errors='coerce').astype('float') # 转换整数型,用Int64(带大写I)支持空值的整数类型 df['sibsp'] = pd.to_numeric(df['sibsp'], errors='coerce').astype('Int64') df['pclass'] = pd.to_numeric(df['pclass'], errors='coerce').astype('Int64') df['survived'] = pd.to_numeric(df['survived'], errors='coerce').astype('Int64')
3. 将空值替换为None
把pandas里的NaN替换成None:
df = df.where(pd.notna(df), None)
你的代码为什么没生效?
你定义的new_type是空列表,后续的last、age等变量只是单个值的转换,既没加入列表,也没遍历处理所有数据行,自然不会作用到输出的二维列表上。如果一定要用原生列表处理(不推荐,效率低),可以这样写:
# 假设你的原始数据是data_list(即输出的二维列表) data_list = [['Braund', ' Mr. Owen Harris', 'male', '22', '1', '3', '7.25', '0'], ['Cumings', ' Mrs. John Bradley (Florence Briggs Thayer)', 'female', '38', '1', '1', '71.2833', '1'], ['Heikkinen', ' Miss. Laina', 'female', '26', '0', '3', '7.925', '1']] # 按列顺序定义转换函数,空值转None converters = [ str, str, str, lambda x: float(x) if x.strip() else None, lambda x: int(x) if x.strip() else None, lambda x: int(x) if x.strip() else None, lambda x: float(x) if x.strip() else None, lambda x: int(x) if x.strip() else None ] processed_data = [] for row in data_list: processed_row = [] for val, func in zip(row, converters): try: processed_row.append(func(val)) except: processed_row.append(None) # 转换失败也设为None processed_data.append(processed_row)
内容的提问来源于stack exchange,提问作者Dhruvina Desai
相关产品推荐
相关产品推荐

