如何将Numpy数组转换为忽略0值的样本簇频率数组?
Solution to Compute Cluster Frequency Array in NumPy
Hey there! Let's work through how to generate that 5-dimensional frequency array you need. The goal is to calculate, for each sample row, the frequency of clusters 1-5 (ignoring 0s), with the sum of frequencies per row equal to 1.
Step-by-Step Approach
- Count occurrences of each cluster (1-5) per row: We'll use NumPy's
bincountto tally how many times each cluster appears in every row. - Calculate non-zero element count per row: This gives us the denominator for our frequency calculation (since we ignore 0s).
- Compute frequencies: Divide each cluster's count by the non-zero count of its row to get normalized frequencies.
- Format output: Match the example's structure with column headers and clean formatting.
Code Implementation
import numpy as np # Your original array (replace with your input) array = np.random.randint(6, size=(50, 400)) # Example input from your question (uncomment to test) # array = np.array([[0, 4, 2, 3, 0], # [5, 5, 2, 5, 3]]) # Step 1: Count occurrences of clusters 1-5 for each row # np.bincount returns counts for 0-5, we slice [1:] to get 1-5 row_counts = np.array([np.bincount(row, minlength=6)[1:] for row in array]) # Step 2: Get number of non-zero elements per row (keepdims for broadcasting) non_zero_total = np.count_nonzero(array, axis=1, keepdims=True) # Step 3: Compute frequencies, handle rows with all zeros (avoid division by zero) freq_array = np.where(non_zero_total == 0, 0, row_counts / non_zero_total) # Step 4: Print formatted output print("1 2 3 4 5") for freq_row in freq_array: # Format values: show 0 instead of 0.00, and 2 decimal places for others formatted = ['0' if val == 0 else f"{val:.2f}" for val in freq_row] print(' '.join(formatted))
Explanation
np.bincount(row, minlength=6): This function counts the frequency of each integer in the row. Settingminlength=6ensures we get counts for 0 through 5 even if some values don't appear. We slice[1:]to discard the count of 0s and keep only clusters 1-5.np.count_nonzero(..., keepdims=True): Counts non-zero elements per row.keepdims=Truekeeps the result as a 2D array, which lets us use NumPy's broadcasting to divide each row's counts by its non-zero total.np.where(...): Handles edge cases where a row has all zeros (avoids division by zero by setting frequencies to 0 for those rows).- Formatting: The output loop converts values to strings, showing "0" instead of "0.00" for cleaner readability, matching your example.
内容的提问来源于stack exchange,提问作者Jack Arnestad
相关产品推荐
相关产品推荐

