You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

HousePrices数据集:如何将city等object类型转为String类型?

解决方法

1. 正确将Object类型列转为String类型

针对city、country、street、statezip这类字段,若你的pandas版本≥1.0,用以下两种方式都能完成转换:

  • 批量转换最简写法:
cols_to_convert = ['city', 'country', 'street', 'statezip']
df[cols_to_convert] = df[cols_to_convert].astype('string')
  • 更严谨的写法(支持专门的缺失值标识):
from pandas import StringDtype
df[cols_to_convert] = df[cols_to_convert].astype(StringDtype())

要是转换失败,先排查列里的混合类型问题——比如有些单元格可能是数字、None而非字符串,先统一转成字符串再转String类型:

df[cols_to_convert] = df[cols_to_convert].astype(str).astype('string')

还可以先查看列的元素类型分布确认问题:

for col in cols_to_convert:
    print(f"列 {col} 的元素类型分布:")
    print(df[col].apply(type).value_counts())

2. 修复groupby计算mean的错误

groupby计算mean报错,大概率不是字符串列的锅——pandas默认会忽略非数值列的聚合操作。问题通常出在你指定了非数值列进行mean计算,或者分组逻辑有误。正确操作如下:

  • 仅对数值列(比如price、sqft_living)做均值聚合:
# 按city分组,计算房价均值
df.groupby('city')['price'].mean()

# 多列均值计算
df.groupby('city')[['price', 'sqft_living']].mean()

如果需要保留分组内的字符串信息,同时聚合数值列,用agg指定每个列的处理方式:

df.groupby('city').agg({
    'price': 'mean',
    'country': 'first'  # 取分组内第一个country值(假设同城市country一致)
})

3. 内存优化验证

转成String类型后内存占用会比原Object类型更低,你可以用下面的命令验证:

print("转换前目标列内存总和:", df[cols_to_convert].memory_usage(deep=True).sum())
print("转换后目标列内存总和:", df[cols_to_convert].memory_usage(deep=True).sum())

内容的提问来源于stack exchange,提问作者PradM

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 21:12:39