生成可导入Pandas DataFrame的元组列表:求解单词最大辅音簇长度
Got it, let's fix this step by step. You're already extracting consonant clusters correctly—now we just need to calculate the maximum length for each word's clusters and package that into tuples ready for Pandas.
Here's a complete, working solution:
import re import pandas as pd list_2 = ['financial','disastrous','accuracy','important','numbers'] # Note: Your regex doesn't include 'c'—if you want to capture clusters like 'cc' in 'accuracy', add 'c' to the character set reg = r'[bdðfghjklmnprstvxþ]+' # Build the tuple list for Pandas cluster_lengths = [] for word in list_2: # Extract consonant clusters for the current word clusters = re.findall(reg, word, re.IGNORECASE) # Calculate the longest cluster length (handle empty case for safety) max_cluster_len = max(len(cluster) for cluster in clusters) if clusters else 0 # Add (word, max length) tuple to our list cluster_lengths.append((word, max_cluster_len)) # Create the DataFrame df = pd.DataFrame(cluster_lengths, columns=['Word', 'Max Consonant Cluster Length']) # Verify the results print("Tuple list for DataFrame:", cluster_lengths) print("\nPandas DataFrame:") print(df)
Output:
The cluster_lengths tuple list will be:[('financial', 1), ('disastrous', 3), ('accuracy', 1), ('important', 2), ('numbers', 2)]
The resulting DataFrame will look like this:
| Word | Max Consonant Cluster Length |
|---|---|
| financial | 1 |
| disastrous | 3 |
| accuracy | 1 |
| important | 2 |
| numbers | 2 |
Key Details:
- We use a generator expression
max(len(cluster) for cluster in clusters)to efficiently compute the longest cluster length for each word. Theif clusters else 0handles edge cases where a word might have no consonants (though none exist in your sample list). - The tuple list is directly compatible with Pandas—we just specify column names when initializing the DataFrame.
- Quick side note: Your regex is missing the consonant 'c'. If you intended to capture clusters like the 'cc' in 'accuracy', update the regex to
r'[bcðfghjklmnprstvxþ]+'(add 'c' to the character set). This would change the max length for 'accuracy' from 1 to 2.
内容的提问来源于stack exchange,提问作者Paul Richardson
相关产品推荐
相关产品推荐

