为何Category类型列占用空间比Object类型列更大?
Category类型DataFrame占用内存反而更大的原因
运行以下代码并查看info()输出后发现,使用Category类型的DataFrame占用932字节空间,反而比使用Object类型的DataFrame(624字节)占用空间更大。
测试代码
def initData(): myPets = {"animal": ["cat", "alligator", "snake", "dog", "gerbil", "lion", "gecko", "hippopotamus", "parrot", "crocodile", "falcon", "hamster", "guinea pig"], "feel" : ["furry", "rough", "scaly", "furry", "furry", "furry", "rough", "rough", "feathery", "rough", "feathery", "furry", "furry" ], "where lives": ["indoor", "outdoor", "indoor", "indoor", "indoor", "outdoor", "indoor", "outdoor", "indoor", "outdoor", "outdoor", "indoor", "indoor" ], "risk": ["safe", "dangerous", "dangerous", "safe", "safe", "dangerous", "safe", "dangerous", "safe", "dangerous", "safe", "safe", "safe" ], "favorite food": ["treats", "fish", "bugs", "treats", "grain", "antelope", "bugs", "antelope", "grain", "fish", "rabbit", "grain", "grain" ], "want to own": [1, 0, 0, 1, 1, 0, 1, 0, 1, 0, 1, 1, 1 ] } petDF = pd.DataFrame(myPets) petDF = petDF.set_index("animal") #print(petDF.info()) #petDF.head(100) return petDF def addCategoryColumns(myDF): myDF["cat_feel"] = myDF["feel"].astype("category") myDF["cat_where_lives"] = myDF["where lives"].astype("category") myDF["cat_risk"] = myDF["risk"].astype("category") myDF["cat_favorite_food"] = myDF["favorite food"].astype("category") return myDF objectsDF = initData() categoriesDF = initData() categoriesDF = addCategoryColumns(categoriesDF) categoriesDF = categoriesDF.drop(["feel", "where lives", "risk", "favorite food"], axis = 1) print(objectsDF.info()) print(categoriesDF.info()) categoriesDF.head()
输出结果
<class 'pandas.core.frame.DataFrame'> Index: 13 entries, cat to guinea pig Data columns (total 5 columns): # Column Non-Null Count Dtype --- ------ -------------- ----- 0 feel 13 non-null object 1 where lives 13 non-null object 2 risk 13 non-null object 3 favorite food 13 non-null object 4 want to own 13 non-null int64 dtypes: int64(1), object(4) memory usage: 624.0+ bytes None <class 'pandas.core.frame.DataFrame'> Index: 13 entries, cat to guinea pig Data columns (total 5 columns): # Column Non-Null Count Dtype --- ------ -------------- ----- 0 want to own 13 non-null int64 1 cat_feel 13 non-null category 2 cat_where_lives 13 non-null category 3 cat_risk 13 non-null category 4 cat_favorite_food 13 non-null category dtypes: category(4), int64(1) memory usage: 932.0+ bytes None
原因解析
这是因为Category类型存在固定的元数据开销:
- 每个Category列都需要额外存储该列的类别列表(比如
feel列需要保存furry、rough、scaly、feathery这几个类别字符串),这部分会占用额外内存 - 当数据集规模极小(这里仅13行)时,用整数编码代替字符串节省的空间,远不足以抵消类别元数据的开销
- 只有当数据量足够大,且列的不同类别占比极低(比如某列只有2个类别,但有上万行数据)时,Category类型的内存优势才会明显体现出来
内容的提问来源于stack exchange,提问作者nicomp
相关产品推荐
相关产品推荐

