You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

统计给定ArrayList中长度>3的字符串组合的出现次数

解决双词组合统计问题:仅统计由长度>3的单词组成的连续词对

Hey folks, let's tackle this problem step by step. The core requirement here is to count how often consecutive two-word pairs show up in a given text—but only if both words in the pair are longer than 3 characters. Any pair that includes a word with 3 or fewer characters gets tossed out, like the "they are" or "are successful" pairs in the example.

Breakdown of the Approach

Here's how we can build a practical solution (using Python as an example):

  1. Clean up the text: First, strip out punctuation, normalize capitalization, and split the text into individual words. This ensures we don't count "scientists" and "scientists," as separate entries.
  2. Filter valid pairs: Loop through the list of words, grab each consecutive pair, and check if both words are longer than 3 characters. Only keep the pairs that meet this rule.
  3. Count occurrences: Use a counter tool to tally up how many times each valid pair appears.
  4. Format the output: Convert the counter results into the requested format (e.g., scientists found-2).

Working Code Example

from collections import Counter
import re

# The input text we're analyzing
input_text = "scientists found way to reduce global warming scientists, found way to minimize water pollution scientists said that they are successful Rony said that they are successful johnny said that he failed"

# Step 1: Clean and split the text
# Remove punctuation, convert to lowercase, split into individual words
cleaned_words = re.sub(r'[^\w\s]', '', input_text).lower().split()

# Step 2: Generate valid two-word pairs
valid_pairs = []
for i in range(len(cleaned_words) - 1):
    first_word = cleaned_words[i]
    second_word = cleaned_words[i + 1]
    # Only keep pairs where both words are longer than 3 characters
    if len(first_word) > 3 and len(second_word) > 3:
        valid_pairs.append(f"{first_word} {second_word}")

# Step 3: Count the frequency of each valid pair
pair_counts = Counter(valid_pairs)

# Step 4: Format and print the final result
final_output = ' '.join([f"{pair}-{count}" for pair, count in pair_counts.items()])
print(final_output)

Output Result

Running this code will produce exactly the example output we need:

scientists found-2 said that-3

Key Details to Note

  • Punctuation handling: The regex re.sub(r'[^\w\s]', '', input_text) takes care of commas, periods, and other stray punctuation that might attach to words.
  • Case insensitivity: Converting everything to lowercase ensures "Scientists" and "scientists" are treated as the same word.
  • Strict length check: We only retain pairs where both words are longer than 3 characters—so even if one word is valid but the other isn't, the pair gets excluded.

内容的提问来源于stack exchange,提问作者Vikki

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 04:07:57