使用Numpy从二维数组生成连续值转换频次汇总DataFrame
Hey there! Let's walk through how to build that transition count DataFrame you need. Here's a step-by-step solution using NumPy and Pandas, which fits perfectly with your existing array setup.
Step 1: Set up your imports and existing array
First, let's start with the tools we need and your provided array (included for completeness):
import numpy as np import pandas as pd from collections import Counter # Your given 20x10 array arr = np.array([[0, 9, 9, 7, 6, 2, 6, 4, 4, 3], [0, 2, 1, 7, 1, 0, 2, 6, 6, 2], [7, 3, 9, 8, 9, 7, 1, 10, 4, 2], [0, 7, 0, 1, 4, 5, 8, 4, 2, 2], [5, 2, 12, 3, 12, 2, 7, 12, 4, 12], [0, 11, 0, 10, 7, 4, 12, 11, 11, 4], [0, 9, 9, 8, 5, 11, 7, 6, 10, 7], [0, 9, 0, 10, 11, 1, 5, 10, 8, 10], [3, 11, 4, 7, 7, 8, 10, 11, 5, 12], [0, 5, 0, 8, 1, 5, 1, 11, 9, 1], [0, 8, 6, 12, 11, 1, 4, 11, 4, 1], [2, 10, 5, 5, 7, 9, 11, 6, 12, 10], [9, 8, 11, 4, 10, 1, 10, 12, 0, 3], [0, 7, 10, 8, 2, 10, 5, 7, 9, 6], [0, 9, 6, 9, 1, 12, 4, 1, 8, 2], [8, 12, 10, 12, 8, 2, 3, 0, 11, 4], [6, 7, 11, 12, 8, 7, 1, 9, 9, 8], [0, 4, 0, 8, 9, 7, 1, 1, 3, 5], [0, 8, 1, 11, 2, 12, 6, 11, 12, 10], [0, 7, 3, 8, 3, 3, 7, 1, 9, 9]])
Step 2: Extract all consecutive value pairs
We need to pull every instance where one value is immediately followed by another in each row. For example, in the first row [0,9,9,...], we get pairs like (0,9), (9,9), etc.:
# Collect all consecutive value pairs from every row transition_pairs = [] for row in arr: # Pair each element with the one right after it transition_pairs.extend(zip(row[:-1], row[1:]))
Step 3: Count occurrences of each transition
Use Counter to tally how many times each (from_value, to_value) pair shows up:
# Count how often each transition happens pair_counts = Counter(transition_pairs)
Step 4: Create and populate the transition DataFrame
Make an empty DataFrame with rows and columns indexed 0-12 (all starting at 0), then fill it with our counts:
# Initialize empty DataFrame with 0-12 as row/column labels transition_df = pd.DataFrame(0, index=np.arange(0,13), columns=np.arange(0,13)) # Fill the DataFrame with transition counts for (from_val, to_val), count in pair_counts.items(): transition_df.loc[from_val, to_val] = count
Verify the results
Let's check your examples to confirm it works:
transition_df.loc[0, 9]returns 4 (matches your stated count of 0→9 transitions)transition_df.loc[10, 12]returns 2 (matches your 10→12 count)
The final DataFrame will have every cell (x,y) showing how many times value x was immediately followed by value y across all rows in your original array.
内容的提问来源于stack exchange,提问作者Abhishek Jain

