统计给定ArrayList中长度>3的字符串组合的出现次数
解决双词组合统计问题:仅统计由长度>3的单词组成的连续词对
Hey folks, let's tackle this problem step by step. The core requirement here is to count how often consecutive two-word pairs show up in a given text—but only if both words in the pair are longer than 3 characters. Any pair that includes a word with 3 or fewer characters gets tossed out, like the "they are" or "are successful" pairs in the example.
Breakdown of the Approach
Here's how we can build a practical solution (using Python as an example):
- Clean up the text: First, strip out punctuation, normalize capitalization, and split the text into individual words. This ensures we don't count "scientists" and "scientists," as separate entries.
- Filter valid pairs: Loop through the list of words, grab each consecutive pair, and check if both words are longer than 3 characters. Only keep the pairs that meet this rule.
- Count occurrences: Use a counter tool to tally up how many times each valid pair appears.
- Format the output: Convert the counter results into the requested format (e.g.,
scientists found-2).
Working Code Example
from collections import Counter import re # The input text we're analyzing input_text = "scientists found way to reduce global warming scientists, found way to minimize water pollution scientists said that they are successful Rony said that they are successful johnny said that he failed" # Step 1: Clean and split the text # Remove punctuation, convert to lowercase, split into individual words cleaned_words = re.sub(r'[^\w\s]', '', input_text).lower().split() # Step 2: Generate valid two-word pairs valid_pairs = [] for i in range(len(cleaned_words) - 1): first_word = cleaned_words[i] second_word = cleaned_words[i + 1] # Only keep pairs where both words are longer than 3 characters if len(first_word) > 3 and len(second_word) > 3: valid_pairs.append(f"{first_word} {second_word}") # Step 3: Count the frequency of each valid pair pair_counts = Counter(valid_pairs) # Step 4: Format and print the final result final_output = ' '.join([f"{pair}-{count}" for pair, count in pair_counts.items()]) print(final_output)
Output Result
Running this code will produce exactly the example output we need:
scientists found-2 said that-3
Key Details to Note
- Punctuation handling: The regex
re.sub(r'[^\w\s]', '', input_text)takes care of commas, periods, and other stray punctuation that might attach to words. - Case insensitivity: Converting everything to lowercase ensures "Scientists" and "scientists" are treated as the same word.
- Strict length check: We only retain pairs where both words are longer than 3 characters—so even if one word is valid but the other isn't, the pair gets excluded.
内容的提问来源于stack exchange,提问作者Vikki
相关产品推荐
相关产品推荐

