Neo4j属性内存缓存机制及索引查询性能疑问
Great questions—these get right to how Neo4j balances disk-based persistence with in-memory speed, which is key to its performance. Let’s break them down one by one:
Neo4j doesn’t cache properties in isolation—it uses a page-centric caching model centered around its core Page Cache. Here’s the breakdown:
- All data (nodes, relationships, properties, indexes) is stored on disk in fixed-size blocks called pages (usually 8KB by default).
- When you access a node or relationship’s properties, Neo4j first checks if the entire page containing that property record is already in the Page Cache. If not, it loads the entire page from disk into memory.
- Once the page is in the cache, any subsequent access to properties (or other records) on that same page happens at in-memory speed.
- The Page Cache uses an LRU (Least Recently Used) eviction policy, so it prioritizes keeping frequently accessed pages in memory while flushing less-used ones back to disk when space is tight.
It’s worth noting that properties are stored alongside node/relationship metadata in most cases (unless you have extremely large properties, which get stored as separate "overflow" pages—but those are still cached the same way when accessed).
First, a quick correction: indexes in Neo4j are also stored as disk pages, just like node/relationship data. When you run an indexed query, here’s what happens step by step:
- Index Page Loading: If the relevant index pages (Neo4j uses B-tree indexes for property lookups) aren’t in the Page Cache, Neo4j will load them from disk. B-trees are designed to minimize disk I/O—most indexes only have 3-4 levels, so this means just 3-4 disk reads to traverse the index and find the pointer to your target node/relationship.
- Data Page Loading: Once the index gives the location of the target node/relationship’s page, Neo4j loads that data page into the Page Cache (if it’s not already there).
- In-Memory Access: Subsequent queries for the same indexed property or node will hit the cached index and data pages, resulting in near-instant access.
Even when indexes aren’t initially in memory, the lookup is still way faster than a full scan (which would require reading every page in the database). The O(log n) time complexity still holds here—each step of the B-tree traversal is a disk read, but log n is a small number even for large datasets (e.g., log₂(100 million) is ~27, but in practice Neo4j’s B-trees are wider, so the number of disk reads is even lower, like 3-4).
Over time, the Page Cache will keep frequently accessed index and data pages in memory, so your indexed queries will get faster as the cache warms up.
内容的提问来源于stack exchange,提问作者CypherFancy

