如何使用Python拆分连续表情符号?实现单个表情独立拆分
Got it, let's solve this problem where you want to split a string so that consecutive emojis become individual elements instead of being grouped together. Here's how you can do it in Python, and it works with all Unicode emojis:
Approach 1: Split First, Then Process Emoji Blocks
This method first splits the string by whitespace, then breaks down any blocks of consecutive emojis into individual characters:
import re # Regex pattern to match most Unicode emojis emoji_pattern = re.compile( "[" u"\U0001F600-\U0001F64F" # Emoticons u"\U0001F300-\U0001F5FF" # Symbols & Pictographs u"\U0001F680-\U0001F6FF" # Transport & Map Symbols u"\U0001F1E0-\U0001F1FF" # Flags (iOS) u"\U00002500-\U00002BEF" # Misc symbols u"\U00002702-\U000027B0" u"\U000024C2-\U0001F251" u"\U0001F926-\U0001F937" u"\U00010000-\U0010FFFF" u"\u2640-\u2642" u"\u2600-\u2B55" u"\u200d" u"\u23CF" u"\u23E9" u"\u231A" u"\ufe0f" # Variation Selector-16 u"\u3030" "]+", flags=re.UNICODE ) def split_with_emojis(input_str): # Split by whitespace first initial_split = input_str.split() final_result = [] for item in initial_split: # Check if the item is entirely made of emojis if emoji_pattern.fullmatch(item): # Split the emoji block into individual emojis final_result.extend(list(item)) else: # Add non-emoji items as-is final_result.append(item) return final_result # Test it out EM = 'Hey 😷😷😷' print(split_with_emojis(EM)) # Output: ['Hey', '😷', '😷', '😷']
How This Works:
- The regex pattern covers nearly all standard Unicode emojis, so it'll recognize everything from smileys to flags to special symbols.
- We first split the string normally by whitespace, then check each segment: if it's a full block of emojis, we convert it to a list (which splits each emoji into its own element) and add those to the result. Non-emoji words stay intact.
Approach 2: Single Regex Findall (Shorter)
This uses a regex to directly match either full words or individual emojis, then extracts all matches in one step:
import re # Regex to match either words (letters/numbers) or single emojis combined_pattern = re.compile( r'\w+|' # Match words (a-z, A-Z, 0-9, underscores) r'[\U0001F600-\U0001F64F\U0001F300-\U0001F5FF\U0001F680-\U0001F6FF\U0001F1E0-\U0001F1FF\U00002500-\U00002BEF\U00002702-\U000027B0\U000024C2-\U0001F251\U0001F926-\U0001F937\U00010000-\U0010FFFF\u2640-\u2642\u2600-\u2B55\u200d\u23CF\u23E9\u231A\ufe0f\u3030]' ) EM = 'Hey 😷😷😷' result = combined_pattern.findall(EM) print(result) # Output: ['Hey', '😷', '😷', '😷']
Note:
If you need to handle non-word characters (like punctuation attached to words, e.g., 'Hey!'), adjust the \w+ part of the regex. For example, use r'[^\s]+' instead to match any sequence of non-whitespace characters, then add logic to split if needed—but for most cases where words are separated by spaces, \w+ works fine.
内容的提问来源于stack exchange,提问作者Abi2021

