You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中预计算三维键值数据的高效访问最优方案(附代码示例)

Hey there! Based on your needs—super fast access to precomputed 3D sorted data (with keys h1, h2, q mapping to n and p)—here are a few optimized approaches that fully leverage your data's sorted nature, way better than nested dicts or SQLite:

1. Custom Integer Index Mapping (O(1) Access, Fastest Option)

Since your keys are sorted and have discrete values, you can map each unique h1, h2, q to an integer index, then store your data in a 3D NumPy array. This gives you instant O(1) access because array indexing is direct memory lookup.

Step-by-Step Implementation:

First, extract and sort all unique keys from your existing dictionary:

import numpy as np

# Your original data dict
dic = {0.6: {0.6: {25: [0.721015194906449, 38.5765797217578], 30: [0.894997480537817, 41.4758084379593], 35: [1.09507774740190, 44.4203224961566], 40: [1.32615185231568, 47.4085304722462], 45: [1.59288184615660, 50.4387051155152]}, 0.8: {25: [0.699928591599505, 38.1952586298542], 30: [0.870503808316267, 41.0984549056602], 35: [1.06723254390729, 44.0475150243940], 40: [1.29499245972316, 47.0407802928774], 45: [1.55841136584623, 50.0764537801238]}, 1.0: {25: [0.678904933323108, 37.8080339255472], 30: [0.846078073485272, 40.7153829050255], 35: [1.03946048887268, 43.6691963099343], 40: [1.26390856173115, 46.6677433191870], 45: [1.52401375488507, 49.7091525400001]}}, 0.8: {0.6: {25: [0.699928591599505, 38.1952586298542], 30: [0.870503808316267, 41.0984549056602], 35: [1.06723254390729, 44.0475150243940], 40: [1.29499245972316, 47.0407802928774], 45: [1.55841136584623, 50.0764537801238]}, 0.8: {25: [0.678904933323108, 37.8080339255472], 30: [0.846078073485272, 40.7153829050255], 35: [1.03946048887268, 43.6691963099343], 40: [1.26390856173115, 46.6677433191870], 45: [1.52401375488507, 49.7091525400001]}, 1.0: {25: [0.657945584308118, 37.4146226119758], 30: [0.821721975385228, 40.3263243022160], 35: [1.01176326951241, 43.2851144574096], 40: [1.23290137939785, 46.2891847659982], 45: [1.48968930734260, 49.3365841236722]}}, 1.0: {0.6: {25: [0.678904933323108, 37.8080339255472], 30: [0.846078073485272, 40.7153829050255], 35: [1.03946048887268, 43.6691963099343], 40: [1.26390856173115, 46.6677433191870], 45: [1.52401375488507, 49.7091525400001]}, 0.8: {25: [0.657945584308118, 37.4146226119758], 30: [0.821721975385228, 40.3263243022160], 35: [1.01176326951241, 43.2851144574096], 40: [1.23290137939785, 46.2891847659982], 45: [1.48968930734260, 49.3365841236722]}, 1.0: {25: [0.637052017659749, 37.0147183330305], 30: [0.797437318766716, 39.9309893301663], 35: [0.984142627694414, 42.8949977775162], 40: [1.20197209243436, 45.9048519379388], 45: [1.45543814550046, 48.9585152177077]}}}

# Extract sorted unique keys
sorted_h1 = sorted(dic.keys())
sorted_h2 = sorted(next(iter(dic.values())).keys())
sorted_q = sorted(next(iter(next(iter(dic.values())).values())).keys())

# Create index mappings (float -> integer)
h1_to_idx = {h: i for i, h in enumerate(sorted_h1)}
h2_to_idx = {h: i for i, h in enumerate(sorted_h2)}
q_to_idx = {q: i for i, q in enumerate(sorted_q)}

Then build a 3D array to store n and p:

# Initialize 3D array: (len(h1), len(h2), len(q), 2) where [:, :, :, 0] is n, [:, :, :, 1] is p
data_array = np.zeros((len(sorted_h1), len(sorted_h2), len(sorted_q), 2), dtype=np.float64)

# Populate the array
for h1_idx, h1 in enumerate(sorted_h1):
    for h2_idx, h2 in enumerate(sorted_h2):
        for q_idx, q in enumerate(sorted_q):
            n_val, p_val = dic[h1][h2][q]
            data_array[h1_idx, h2_idx, q_idx] = [n_val, p_val]

Finally, write your desired access functions:

def n(h1, h2, q):
    # Handle float precision by rounding to match your data's decimal places
    idx_h1 = h1_to_idx.get(round(h1, 1))
    idx_h2 = h2_to_idx.get(round(h2, 1))
    idx_q = q_to_idx.get(q)
    if None in (idx_h1, idx_h2, idx_q):
        raise ValueError("Key not found in precomputed data")
    return data_array[idx_h1, idx_h2, idx_q, 0]

def p(h1, h2, q):
    idx_h1 = h1_to_idx.get(round(h1, 1))
    idx_h2 = h2_to_idx.get(round(h2, 1))
    idx_q = q_to_idx.get(q)
    if None in (idx_h1, idx_h2, idx_q):
        raise ValueError("Key not found in precomputed data")
    return data_array[idx_h1, idx_h2, idx_q, 1]

Why this works:

  • No hash lookups or search overhead—just direct memory access.
  • NumPy arrays are memory-efficient, way more so than nested dictionaries.
  • Float key precision issues are easily handled by rounding to the decimal places used in your data.

2. NumPy Structured Array with Binary Search (Great for Sorted Data)

If you prefer a flat structure but still want to leverage sorted keys, you can use a structured array and np.searchsorted to find matches quickly.

Implementation:

# Flatten your data into a list of tuples
flat_data = []
for h1, h2_dict in dic.items():
    for h2, q_dict in h2_dict.items():
        for q, (n_val, p_val) in q_dict.items():
            flat_data.append((h1, h2, q, n_val, p_val))

# Convert to structured array, sorted by h1 -> h2 -> q
structured_arr = np.array(flat_data, dtype=[('h1', 'f8'), ('h2', 'f8'), ('q', 'i4'), ('n', 'f8'), ('p', 'f8')])
structured_arr.sort(order=['h1', 'h2', 'q'])

# Helper function to find indices via binary search
def get_indices(h1, h2, q):
    # Narrow down by h1 first
    h1_mask = structured_arr['h1'] == h1
    h1_subset = structured_arr[h1_mask]
    # Then narrow down by h2
    h2_mask = h1_subset['h2'] == h2
    h2_subset = h1_subset[h2_mask]
    # Final binary search for q
    q_idx = np.searchsorted(h2_subset['q'], q)
    if q_idx < len(h2_subset) and h2_subset['q'][q_idx] == q:
        return h2_subset[q_idx]
    raise ValueError("Key not found")

# Your access functions
def n(h1, h2, q):
    return get_indices(h1, h2, q)['n']

def p(h1, h2, q):
    return get_indices(h1, h2, q)['p']

Pros/Cons:

  • Ideal if you need to run additional operations on the data (like filtering or slicing).
  • Slightly slower than the index mapping approach but still way faster than dicts/SQLite.

3. Memory-Mapped NumPy Array (For Very Large Datasets)

If your data is too big to fit in RAM, use np.memmap to store the array on disk while accessing it like a regular NumPy array. This keeps memory usage low while maintaining fast access speeds.

Quick Example:

# Create and save the memmap array (run once during precomputation)
mmap_array = np.memmap('precomputed_data.npy', dtype=np.float64, mode='w+', shape=(len(sorted_h1), len(sorted_h2), len(sorted_q), 2))
# Populate it using the same loop as option 1
# ...
mmap_array.flush()  # Save changes to disk

# Load in your main program
mmap_array = np.memmap('precomputed_data.npy', dtype=np.float64, mode='r', shape=(len(sorted_h1), len(sorted_h2), len(sorted_q), 2))

# Use the same index mappings and access functions as option 1

Why this beats SQLite:

  • Disk access uses contiguous blocks (far faster than SQLite's row-based lookups).
  • No SQL query parsing overhead—direct array access.

Why Your Previous Approaches Were Slow:

  • Nested Dicts: Even with splitting, dict lookups involve hash table checks, and nested structures add unnecessary overhead. Dicts also have much higher memory usage compared to arrays.
  • SQLite: Disk-based queries have inherent IO latency, plus the overhead of parsing SQL statements and row-by-row retrieval—great for complex queries, but overkill for simple key-based lookups.

内容的提问来源于stack exchange,提问作者rhea

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 10:22:34