You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

《Database System Concepts》6th Ed.主索引处理非键相等查询连续读取疑问

为什么用主索引处理非键属性等值查询时读取是连续的?

Great question—this is a common point of confusion when first learning about clustered (primary) indexes. Let’s break this down step by step using the context from Database System Concepts 6th ed.

First, let’s recap what a primary (clustered) index entails:

  • The entire data file is physically sorted and stored in the order of the clustering key (primary key). So records with adjacent primary key values are stored next to each other on disk.
  • The primary index itself is a sorted structure (like a B-tree) that maps primary key values to the disk blocks where those records reside.

Now, the algorithm you’re referring to—using the primary index to handle an equality constraint on a non-key attribute—doesn’t mean we’re directly using the index to look up the non-key value. Instead, here’s what’s happening:

Since the file is sorted by the search key (primary key), reading the file sequentially allows us to check every record’s non-key attribute for the equality condition.

Wait, that might be the key point you’re missing! The book isn’t saying we’re using the index to jump directly to matching records (we can’t, since the index is keyed on the primary key, not the non-key attribute). Instead, the algorithm is leveraging the fact that the data file is physically contiguous and sorted by the primary key, so reading the entire file sequentially is an efficient operation.

But why is this considered "using the primary index"? Because the primary index’s existence is what enforces the physical sorting of the data file. Without a clustered primary index, the data file would be stored in heap order (unordered), and sequential reads would still work—but the book is highlighting that the sorted order (from the primary index) allows for this sequential read approach to be a valid (and in some cases, the only) way to handle the query when no secondary index exists for the non-key attribute.

Your confusion makes total sense: you’re right that non-key attribute values aren’t sorted, so matching records could be scattered throughout the sorted data file. But the "continuous reading" here refers to how we access the disk blocks—since the data file is stored as a contiguous sorted sequence, we can read blocks one after another without seeking all over the disk. Even though the matching records are scattered within that sequence, the disk I/O is sequential (which is much faster than random I/O), which is what the book is emphasizing.

To put it simply:

  • The data file is physically contiguous and ordered by primary key (thanks to the primary index).
  • When querying a non-key attribute with no secondary index, we have to scan the entire file.
  • Scanning a contiguous, sorted file means we do sequential disk reads (no random jumps), which is what the book calls "continuous" reading.

内容的提问来源于stack exchange,提问作者alex440

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 03:28:42