Pandas DataFrame二进制值列空间优化:category与int8哪种更高效?
Great question—let’s break this down clearly for your 500k-row, thousands-of-columns binary (0s and 1s only) dataset.
How Each DataType Stores Data
Int8
- Every single 0 or 1 takes up 1 byte per cell. No fancy tricks here—just a straight array of bytes. For one 500k-row column, that’s roughly 476 KB. Multiply by your thousands of columns, and you’ve got your base memory footprint.
Category
- Category dtype works by storing two things: a tiny lookup list of unique values (here just
[0, 1]—this takes practically no space) and an integer array where each cell holds the index of its value in the lookup list. - Since we only have 2 unique values, the index array uses the smallest possible integer type (uint8/int8), which also takes 1 byte per cell. The only extra bit is that tiny lookup table, which is negligible next to 500k rows of data.
So Which Is More Space-Efficient?
For your binary data, int8 and category are practically identical in terms of space. The category’s lookup table adds a tiny amount of overhead, but you’d never notice it in a dataset this large. If you’re obsessing over absolute minimal usage, int8 might win by a few bytes per column—but that’s splitting hairs.
Quick Bonus Tip
If you want to squeeze even more space out, you could look into pandas' sparse arrays or the dedicated BooleanDtype, but those come with tradeoffs (like slower operations in some cases). Since you asked specifically about int8 vs category, though, those two are neck-and-neck for your use case.
内容的提问来源于stack exchange,提问作者jma

