Python如何快速读取文件中字典的指定键值?mmap用法求助
Hey there! Storing a dictionary as a raw string is definitely not ideal for random access—you’re right that reading the whole file every time wastes unnecessary time. Let’s go through some practical solutions, including how to use mmap if you want that low-level control:
1. Use shelve (Built-in Key-Value Store)
This is probably the easiest solution since it’s part of Python’s standard library. shelve acts like a persistent dictionary, letting you fetch values by key directly without loading the entire dataset into memory.
Saving the Dictionary:
import shelve # Store your dictionary dict1 = {1: 'a', 2: 'b', 3: 'c', 7: 'g', 9: 'i'} with shelve.open('my_dict.db') as db: db.update(dict1)
Fetching Values in Another Script:
import shelve keys_to_get = [3,7,9] results = {} with shelve.open('my_dict.db', flag='r') as db: for key in keys_to_get: results[key] = db.get(key, None) # Returns None if key doesn't exist print(results) # Output: {3: 'c', 7: 'g', 9: 'i'}
Pros: Super simple, no custom code needed, handles serialization under the hood.
Cons: Not the fastest for extremely large datasets, but perfect for most use cases.
2. Use dbm (Lightweight Built-in Database)
Another standard library option, dbm is a lower-level key-value store. It’s great if you need something faster than shelve and don’t need to store complex Python objects (your integer keys and string values fit perfectly here).
Saving the Dictionary:
import dbm dict1 = {1: 'a', 2: 'b', 3: 'c', 7: 'g', 9: 'i'} with dbm.open('my_dbm.db', 'c') as db: # dbm requires string keys/values, so convert int keys to strings for key, value in dict1.items(): db[str(key)] = value
Fetching Values:
import dbm keys_to_get = [3,7,9] results = {} with dbm.open('my_dbm.db', 'r') as db: for key in keys_to_get: str_key = str(key) results[key] = db.get(str_key, None).decode('utf-8') if str_key in db else None print(results) # Output: {3: 'c', 7: 'g', 9: 'i'}
Pros: Faster than shelve for simple data, minimal overhead.
Cons: Keys/values must be bytes/strings, so you’ll need minor type conversion.
3. Using mmap for Low-Level Random Access
If you’re dealing with an extremely large dictionary and need direct access to disk without loading everything into memory, mmap is a good choice. The catch is you’ll need to build a custom index to track where each key-value pair lives in the file.
Step 1: Save the Dictionary with an Index
First, we’ll store each serialized key-value pair, plus an index that maps keys to their file offset and length. We’ll write the index at the end of the file for easy retrieval.
import pickle import os def save_large_dict(dict_data, filename): index = {} current_offset = 0 with open(filename, 'wb') as f: # Write each key-value pair and track their positions for key, value in dict_data.items(): serialized = pickle.dumps((key, value)) index[key] = (current_offset, len(serialized)) f.write(serialized) current_offset += len(serialized) # Write the index size (as a fixed 4-byte integer) followed by the index itself index_serialized = pickle.dumps(index) index_size = len(index_serialized) f.write(index_size.to_bytes(4, byteorder='big')) f.write(index_serialized)
Step 2: Fetch Values Using mmap
Now we can map the file into memory, read the index, and jump directly to the positions of the keys we need.
import pickle import mmap def fetch_values_mmap(filename, keys): results = {} with open(filename, 'rb') as f: # Map the entire file into memory (no full load into RAM) mm = mmap.mmap(f.fileno(), 0, access=mmap.ACCESS_READ) # Read the index size from the end of the file mm.seek(-4, os.SEEK_END) index_size = int.from_bytes(mm.read(4), byteorder='big') # Read the index itself mm.seek(-4 - index_size, os.SEEK_END) index = pickle.loads(mm.read(index_size)) # Fetch each requested key for key in keys: if key not in index: results[key] = None continue offset, length = index[key] mm.seek(offset) _, value = pickle.loads(mm.read(length)) results[key] = value mm.close() return results # Usage example keys_to_get = [3,7,9] print(fetch_values_mmap('large_dict.dat', keys_to_get)) # Output: {3: 'c', 7: 'g', 9: 'i'}
Pros: Perfect for massive datasets, avoids loading the entire file into memory.
Cons: Requires custom serialization/indexing code—more work than using shelve or dbm.
Final Recommendation
- For most cases, use
shelve—it’s simple, built-in, and handles all the hard parts for you. - If you need faster access for simple data, go with
dbm. - Only use
mmapif you’re dealing with an extremely large dictionary where memory usage is a critical concern.
内容的提问来源于stack exchange,提问作者XiaXuehai

