Python中已有hash(),为何仍需使用id()函数?
id() when we already have hash()? Great question! At first glance, both functions return integers that seem to uniquely identify objects, but they serve entirely different purposes—let’s break this down clearly.
First, let’s recap what id() does (per Python’s official docs):
id(object) returns the "identity" of an object. This is an integer which is guaranteed to be unique and constant for this object during its lifetime. Two objects with non-overlapping lifetimes may have the same id() value. Practically, it acts like a hash function for uniqueness within an object’s lifetime, but it doesn’t have the property of being hard to reverse-engineer like a typical hash.
Core Differences Between id() and hash()
id()is a memory address marker: In CPython (the most common Python implementation),id()directly returns the integer representation of the object’s memory address. Its sole job is to tell you exactly which specific object in memory you’re dealing with—unique for all objects that exist at the same time.hash()is a hash table optimization tool: It’s designed to generate integers that make lookups in dictionaries, sets, and other hash-based structures fast. Hash values prioritize spread (to minimize collisions) over strict uniqueness—different objects can have the same hash value (e.g.,hash(1) == hash(1.0)), and hash values can even vary between different Python processes (for security against collision attacks).
When You Need id() Instead of hash()
Here are the key scenarios where id() is irreplaceable:
Check if two variables reference the same object: The
isoperator in Python actually checks ifid(a) == id(b). This is crucial when you need to distinguish between two objects that have equal values but live in different memory locations:original = {"name": "Alice"} copy = {"name": "Alice"} print(original is copy) # False—two distinct dict objects print(id(original) == id(copy)) # Same result as aboveYou can’t use
hash()here because while both dicts have the same value, they’re mutable objects and can’t be hashed at all.Track object lifecycles: Since
id()stays constant for an object’s entire lifetime, you can use it to debug memory issues (like leaks) by tracking when an object is created and whether it’s ever garbage-collected. For example, you could log theid()of a problematic object to confirm it’s not being retained unintentionally.Assign unique identifiers to mutable objects: Mutable objects can’t be used as keys in dictionaries (since they’re unhashable), but you can use their
id()as a key to associate data with that specific object—even if the object’s content changes later:my_list = [1, 2, 3] object_metadata = {id(my_list): {"created_at": "2024-05-20"}} my_list.append(4) # List content changes, but id() stays the same print(object_metadata[id(my_list)]) # Still retrieves the metadata
Why hash() Can’t Replace id()
- Hash values aren’t strictly unique: Collisions are expected and allowed, which means two different objects can have the same hash.
id()guarantees uniqueness for all objects currently in memory. - Hash doesn’t work on mutable objects: If you’re dealing with lists, dicts, or custom mutable classes,
hash()will throw aTypeError—butid()still works perfectly. - Hash values aren’t consistent across processes: For security, Python randomizes hash values for some types (like strings) between sessions.
id()is tied to the current process’s memory space, so it’s consistent only within that process but reliable for the object’s lifetime.
In short: hash() exists to make hash tables fast, while id() exists to let you pinpoint exactly which object in memory you’re working with. They solve different problems, so Python gives us both!
内容的提问来源于stack exchange,提问作者Nisba

