R语言DataFrame中按类别统计字符成对出现次数
Hey there! Let's walk through how to solve this problem step by step. I'll use a concrete example to make it easy to follow.
First, let's set up a sample DataFrame that matches the scenario you described:
import pandas as pd # Sample DataFrame data = { 'Category': ['A', 'A', 'B', 'B'], 'Col2': [1, 2, 3, 4], 'Col3': ['X', 'Y', 'Z', 'W'], 'Sequence': ['E2E2CC', 'E2C', 'DDCC', 'DDE2'] } df = pd.DataFrame(data)
Step 1: Create a Function to Generate Overlapping Pairs
We need a function that takes a string and returns all overlapping consecutive 2-character pairs. For example, "E2E2CC" will become ['E2', '2E', 'E2', '2C', 'CC'].
def get_overlapping_char_pairs(sequence): # Only generate pairs if the sequence is at least 2 characters long if len(sequence) < 2: return [] # Generate all overlapping 2-character substrings return [sequence[i:i+2] for i in range(len(sequence) - 1)]
Step 2: Apply the Function and Explode the Results
Next, we'll apply this function to the Sequence column, then "explode" the list of pairs into individual rows so we can count them easily.
# Add a new column with the list of pairs df['Overlapping_Pairs'] = df['Sequence'].apply(get_overlapping_char_pairs) # Explode the list into separate rows exploded_df = df.explode('Overlapping_Pairs').dropna(subset=['Overlapping_Pairs'])
Step 3: Group by Category and Count Pair Occurrences
Finally, we group by Category and Overlapping_Pairs, then count how many times each pair appears per category.
# Calculate the count of each pair per category pair_counts = exploded_df.groupby(['Category', 'Overlapping_Pairs']).size().reset_index(name='Count') print(pair_counts)
Expected Output
Running this code will give you a result like this:
| Category | Overlapping_Pairs | Count |
|---|---|---|
| A | E2 | 3 |
| A | 2E | 1 |
| A | 2C | 2 |
| A | CC | 1 |
| B | DD | 2 |
| B | DC | 1 |
| B | CC | 1 |
| B | DE | 1 |
| B | E2 | 1 |
If You Mean Token-Level Pairs (e.g., 2-character tokens)
If your Sequence column is made up of fixed-length tokens (like "E2", "CC"), and you want overlapping pairs of these tokens instead of individual characters, you can adjust the function like this:
def get_overlapping_token_pairs(sequence, token_length=2): # Split the sequence into fixed-length tokens tokens = [sequence[i:i+token_length] for i in range(0, len(sequence), token_length)] # Generate overlapping pairs of tokens if len(tokens) < 2: return [] return [(tokens[i], tokens[i+1]) for i in range(len(tokens) - 1)]
Then repeat steps 2 and 3—this will count pairs like ('E2', 'E2') or ('DD', 'CC') instead of character-level pairs.
内容的提问来源于stack exchange,提问作者rishi

