You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas中对列内相似词组/语句分组并统计频次?

Got it, let's break this down step by step. You want to count multi-word phrases instead of single words, and your example shows you’re interested in both continuous phrases and possibly phrases that span non-consecutive words (like "A123 B234 data" appearing across sentences even with other words in between). Here are two practical implementations using Pandas and Python to cover both scenarios:

1. Count Continuous Multi-Word Phrases (Standard N-Grams)

This is for strict, consecutive word groups (e.g., "test data" or "A123 B234 C345").

Step 1: Setup & Sample Data

First, let's replicate your example data in a Pandas DataFrame:

import pandas as pd
from collections import Counter
import re

# Your sample sentences
data = {
    "text": [
        "A123 B234 C345 test data.",
        "A123 B234 C345 D555 test data.",
        "A123 B234 test data.",
        "A123 B234 C345 more data."
    ]
}
df = pd.DataFrame(data)

Step 2: Clean the Text

We’ll remove punctuation and split sentences into clean word lists:

def clean_text(text):
    # Strip punctuation (keep letters, numbers, and spaces)
    cleaned = re.sub(r'[^\w\s]', '', text)
    # Split into words (handles extra spaces automatically)
    return cleaned.split()

df["cleaned_words"] = df["text"].apply(clean_text)

Step 3: Generate N-Grams (Continuous Phrases)

Write a function to generate all consecutive phrases of varying lengths (adjust n_min and n_max to match your needs):

def generate_ngrams(word_list, n_range):
    ngrams = []
    for n in n_range:
        # Create all consecutive n-word phrases, joined by spaces
        for i in range(len(word_list) - n + 1):
            phrase = ' '.join(word_list[i:i+n])
            ngrams.append(phrase)
    return ngrams

# Generate phrases from 2 to 5 words long (tweak this range as needed)
n_min, n_max = 2, 5
df["phrases"] = df["cleaned_words"].apply(lambda x: generate_ngrams(x, range(n_min, n_max+1)))

Step 4: Count Phrase Frequencies

Flatten all phrases and count their occurrences:

# Flatten the list of phrases from all sentences
all_phrases = [phrase for sublist in df["phrases"] for phrase in sublist]
phrase_counts = Counter(all_phrases)

# Get top 100 most common phrases (like your original single-word approach)
top_phrases = phrase_counts.most_common(100)

# Example output for your sample data
print("Top continuous phrases:")
for phrase, count in top_phrases[:5]:
    print(f"- {phrase}: {count}")

This will output things like:

Top continuous phrases:
- A123 B234: 4
- test data: 3
- A123 B234 C345: 3
- B234 C345: 3
- C345 test data: 1

2. Count Non-Continuous Keyword Combinations

If your example means you want to count sentences that contain a set of keywords (even with other words in between, like "A123 B234" and "data" appearing anywhere in the sentence), use this approach:

# Define the keyword combinations you want to track
target_combinations = [
    ("A123 B234", "data"),
    ("A123 B234", "test", "data"),
    ("A123 B234", "C345", "test", "data")
]

# Count how many sentences contain all keywords in each combination
combination_counts = {}
for combo in target_combinations:
    sentence_matches = 0
    for words in df["cleaned_words"]:
        sentence_str = ' '.join(words)
        # Check if all parts of the combo exist in the sentence
        if all(part in sentence_str for part in combo):
            sentence_matches += 1
    # Format the combo into a readable phrase
    combination_name = ' '.join(combo)
    combination_counts[combination_name] = sentence_matches

# Print results matching your example
print("\nNon-continuous combination counts:")
for phrase, count in combination_counts.items():
    print(f"- {phrase}: {count}")

This will output exactly what you expected (adjusted for your sample data's actual matches):

Non-continuous combination counts:
- A123 B234 data: 4
- A123 B234 test data: 3
- A123 B234 C345 test data: 2

Quick Notes

  • Use the n-gram method if you want natural, consecutive phrases (great for text mining common expressions).
  • Use the keyword combination method if you’re tracking specific terms that don’t need to be consecutive.

内容的提问来源于stack exchange,提问作者jason

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:05:27