You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言DataFrame中按类别统计字符成对出现次数

Solution to Count Overlapping Character Pairs by Category in Pandas

Hey there! Let's walk through how to solve this problem step by step. I'll use a concrete example to make it easy to follow.

First, let's set up a sample DataFrame that matches the scenario you described:

import pandas as pd

# Sample DataFrame
data = {
    'Category': ['A', 'A', 'B', 'B'],
    'Col2': [1, 2, 3, 4],
    'Col3': ['X', 'Y', 'Z', 'W'],
    'Sequence': ['E2E2CC', 'E2C', 'DDCC', 'DDE2']
}
df = pd.DataFrame(data)

Step 1: Create a Function to Generate Overlapping Pairs

We need a function that takes a string and returns all overlapping consecutive 2-character pairs. For example, "E2E2CC" will become ['E2', '2E', 'E2', '2C', 'CC'].

def get_overlapping_char_pairs(sequence):
    # Only generate pairs if the sequence is at least 2 characters long
    if len(sequence) < 2:
        return []
    # Generate all overlapping 2-character substrings
    return [sequence[i:i+2] for i in range(len(sequence) - 1)]

Step 2: Apply the Function and Explode the Results

Next, we'll apply this function to the Sequence column, then "explode" the list of pairs into individual rows so we can count them easily.

# Add a new column with the list of pairs
df['Overlapping_Pairs'] = df['Sequence'].apply(get_overlapping_char_pairs)

# Explode the list into separate rows
exploded_df = df.explode('Overlapping_Pairs').dropna(subset=['Overlapping_Pairs'])

Step 3: Group by Category and Count Pair Occurrences

Finally, we group by Category and Overlapping_Pairs, then count how many times each pair appears per category.

# Calculate the count of each pair per category
pair_counts = exploded_df.groupby(['Category', 'Overlapping_Pairs']).size().reset_index(name='Count')

print(pair_counts)

Expected Output

Running this code will give you a result like this:

CategoryOverlapping_PairsCount
AE23
A2E1
A2C2
ACC1
BDD2
BDC1
BCC1
BDE1
BE21

If You Mean Token-Level Pairs (e.g., 2-character tokens)

If your Sequence column is made up of fixed-length tokens (like "E2", "CC"), and you want overlapping pairs of these tokens instead of individual characters, you can adjust the function like this:

def get_overlapping_token_pairs(sequence, token_length=2):
    # Split the sequence into fixed-length tokens
    tokens = [sequence[i:i+token_length] for i in range(0, len(sequence), token_length)]
    # Generate overlapping pairs of tokens
    if len(tokens) < 2:
        return []
    return [(tokens[i], tokens[i+1]) for i in range(len(tokens) - 1)]

Then repeat steps 2 and 3—this will count pairs like ('E2', 'E2') or ('DD', 'CC') instead of character-level pairs.

内容的提问来源于stack exchange,提问作者rishi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 06:59:06