You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

将Pandas循环统计代码转换为Numpy实现以提升性能的方法问询

Boosting Performance with NumPy for Your Counting Task

Hey there! I get that you're crunched for time and haven't worked with NumPy before—no worries, I'll break this down into simple, actionable steps that skip the deep theory and focus directly on speeding up your code.

The Core Idea

Your original nested loops can drag with large datasets. NumPy shines at vectorized operations (batch processing without manual loops) which will cut down runtime drastically. Here's the plan:

  1. Convert your Pandas columns to NumPy arrays (super straightforward, no new complex concepts).
  2. Precompute all unique (r, s) pairs in cp along with their counts (this replaces scanning cp over and over for each d/k pair).
  3. Generate all the (d, k) combinations you care about (from devs["d"] and a1km).
  4. Match these combinations to the precomputed counts to get your totals.

The NumPy Code

import numpy as np

# Step 1: Convert Pandas columns to NumPy arrays (one-time, easy conversion)
cp_r = cp["r"].to_numpy()
cp_s = cp["s"].to_numpy()
devs_d = devs["d"].to_numpy()
# Extract k values from a1km (assuming a1km is a dictionary; adjust if it's a list)
a1km_ks = np.array(list(a1km.keys()))

# Step 2: Get all unique (r, s) pairs in cp and their occurrence counts
# Stack r and s columns into a 2D array where each row is (r_value, s_value)
cp_pairs = np.column_stack((cp_r, cp_s))
# Fetch unique pairs and how many times each appears
unique_pairs, counts = np.unique(cp_pairs, axis=0, return_counts=True)

# Step 3: Generate all (d, k) combinations you need to count
# Create a grid of every possible d and k pair
d_grid, k_grid = np.meshgrid(devs_d, a1km_ks, indexing="ij")
# Flatten the grid into a list of (d, k) pairs
target_pairs = np.column_stack((d_grid.ravel(), k_grid.ravel()))

# Step 4: Match target pairs to precomputed counts
# For integer values: combine pairs into a single key for fast searching
max_s_val = cp_s.max() + 1  # Offset to avoid overlap between r and s values
unique_keys = unique_pairs[:, 0] * max_s_val + unique_pairs[:, 1]
target_keys = target_pairs[:, 0] * max_s_val + target_pairs[:, 1]

# Find positions of target keys in the unique keys list
indices = np.searchsorted(unique_keys, target_keys)
# Handle pairs that don't exist in cp (set their count to 0)
mask = indices < len(unique_keys)
total_counts = np.zeros(len(target_pairs), dtype=int)
total_counts[mask] = counts[indices[mask]]

# Step 5: Format output as d, k, total
output = np.column_stack((target_pairs[:, 0], target_pairs[:, 1], total_counts))

# Optional: Convert back to a Pandas DataFrame if you need it
# import pandas as pd
# output_df = pd.DataFrame(output, columns=["d", "k", "total"])

Quick Adjustments for String Values

If your d or k are strings instead of integers, replace the key-combining step with this:

# Convert pairs to tuples for easy comparison
unique_pairs_tuples = [tuple(p) for p in unique_pairs]
sorted_unique = sorted(unique_pairs_tuples)

# Look up each target pair in the sorted unique list
indices = [sorted_unique.index(tuple(p)) if tuple(p) in sorted_unique else -1 for p in target_pairs]
total_counts = np.array([counts[i] if i != -1 else 0 for i in indices], dtype=int)

Why This Is Way Faster

Your original code scans cp once for every d/k pair—this adds up fast with large datasets. With NumPy, we scan cp once to precompute all counts, then just look up the values we need. This cuts the time complexity from O(m*n) (m = number of d/k pairs, n = rows in cp) to O(n log n), which is a massive improvement.

内容的提问来源于stack exchange,提问作者Dervin Thunk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 23:22:29