You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Numpy数组转换为忽略0值的样本簇频率数组?

Solution to Compute Cluster Frequency Array in NumPy

Hey there! Let's work through how to generate that 5-dimensional frequency array you need. The goal is to calculate, for each sample row, the frequency of clusters 1-5 (ignoring 0s), with the sum of frequencies per row equal to 1.

Step-by-Step Approach

  1. Count occurrences of each cluster (1-5) per row: We'll use NumPy's bincount to tally how many times each cluster appears in every row.
  2. Calculate non-zero element count per row: This gives us the denominator for our frequency calculation (since we ignore 0s).
  3. Compute frequencies: Divide each cluster's count by the non-zero count of its row to get normalized frequencies.
  4. Format output: Match the example's structure with column headers and clean formatting.

Code Implementation

import numpy as np

# Your original array (replace with your input)
array = np.random.randint(6, size=(50, 400))

# Example input from your question (uncomment to test)
# array = np.array([[0, 4, 2, 3, 0],
#                   [5, 5, 2, 5, 3]])

# Step 1: Count occurrences of clusters 1-5 for each row
# np.bincount returns counts for 0-5, we slice [1:] to get 1-5
row_counts = np.array([np.bincount(row, minlength=6)[1:] for row in array])

# Step 2: Get number of non-zero elements per row (keepdims for broadcasting)
non_zero_total = np.count_nonzero(array, axis=1, keepdims=True)

# Step 3: Compute frequencies, handle rows with all zeros (avoid division by zero)
freq_array = np.where(non_zero_total == 0, 0, row_counts / non_zero_total)

# Step 4: Print formatted output
print("1 2 3 4 5")
for freq_row in freq_array:
    # Format values: show 0 instead of 0.00, and 2 decimal places for others
    formatted = ['0' if val == 0 else f"{val:.2f}" for val in freq_row]
    print(' '.join(formatted))

Explanation

  • np.bincount(row, minlength=6): This function counts the frequency of each integer in the row. Setting minlength=6 ensures we get counts for 0 through 5 even if some values don't appear. We slice [1:] to discard the count of 0s and keep only clusters 1-5.
  • np.count_nonzero(..., keepdims=True): Counts non-zero elements per row. keepdims=True keeps the result as a 2D array, which lets us use NumPy's broadcasting to divide each row's counts by its non-zero total.
  • np.where(...): Handles edge cases where a row has all zeros (avoids division by zero by setting frequencies to 0 for those rows).
  • Formatting: The output loop converts values to strings, showing "0" instead of "0.00" for cleaner readability, matching your example.

内容的提问来源于stack exchange,提问作者Jack Arnestad

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:36:42