将Pandas循环统计代码转换为Numpy实现以提升性能的方法问询
Hey there! I get that you're crunched for time and haven't worked with NumPy before—no worries, I'll break this down into simple, actionable steps that skip the deep theory and focus directly on speeding up your code.
The Core Idea
Your original nested loops can drag with large datasets. NumPy shines at vectorized operations (batch processing without manual loops) which will cut down runtime drastically. Here's the plan:
- Convert your Pandas columns to NumPy arrays (super straightforward, no new complex concepts).
- Precompute all unique
(r, s)pairs incpalong with their counts (this replaces scanningcpover and over for eachd/kpair). - Generate all the
(d, k)combinations you care about (fromdevs["d"]anda1km). - Match these combinations to the precomputed counts to get your totals.
The NumPy Code
import numpy as np # Step 1: Convert Pandas columns to NumPy arrays (one-time, easy conversion) cp_r = cp["r"].to_numpy() cp_s = cp["s"].to_numpy() devs_d = devs["d"].to_numpy() # Extract k values from a1km (assuming a1km is a dictionary; adjust if it's a list) a1km_ks = np.array(list(a1km.keys())) # Step 2: Get all unique (r, s) pairs in cp and their occurrence counts # Stack r and s columns into a 2D array where each row is (r_value, s_value) cp_pairs = np.column_stack((cp_r, cp_s)) # Fetch unique pairs and how many times each appears unique_pairs, counts = np.unique(cp_pairs, axis=0, return_counts=True) # Step 3: Generate all (d, k) combinations you need to count # Create a grid of every possible d and k pair d_grid, k_grid = np.meshgrid(devs_d, a1km_ks, indexing="ij") # Flatten the grid into a list of (d, k) pairs target_pairs = np.column_stack((d_grid.ravel(), k_grid.ravel())) # Step 4: Match target pairs to precomputed counts # For integer values: combine pairs into a single key for fast searching max_s_val = cp_s.max() + 1 # Offset to avoid overlap between r and s values unique_keys = unique_pairs[:, 0] * max_s_val + unique_pairs[:, 1] target_keys = target_pairs[:, 0] * max_s_val + target_pairs[:, 1] # Find positions of target keys in the unique keys list indices = np.searchsorted(unique_keys, target_keys) # Handle pairs that don't exist in cp (set their count to 0) mask = indices < len(unique_keys) total_counts = np.zeros(len(target_pairs), dtype=int) total_counts[mask] = counts[indices[mask]] # Step 5: Format output as d, k, total output = np.column_stack((target_pairs[:, 0], target_pairs[:, 1], total_counts)) # Optional: Convert back to a Pandas DataFrame if you need it # import pandas as pd # output_df = pd.DataFrame(output, columns=["d", "k", "total"])
Quick Adjustments for String Values
If your d or k are strings instead of integers, replace the key-combining step with this:
# Convert pairs to tuples for easy comparison unique_pairs_tuples = [tuple(p) for p in unique_pairs] sorted_unique = sorted(unique_pairs_tuples) # Look up each target pair in the sorted unique list indices = [sorted_unique.index(tuple(p)) if tuple(p) in sorted_unique else -1 for p in target_pairs] total_counts = np.array([counts[i] if i != -1 else 0 for i in indices], dtype=int)
Why This Is Way Faster
Your original code scans cp once for every d/k pair—this adds up fast with large datasets. With NumPy, we scan cp once to precompute all counts, then just look up the values we need. This cuts the time complexity from O(m*n) (m = number of d/k pairs, n = rows in cp) to O(n log n), which is a massive improvement.
内容的提问来源于stack exchange,提问作者Dervin Thunk

