如何在Python中提取所有表情符号并忽略Fitzpatrick修饰符?
I totally get the frustration here—those tiny modifier characters (like skin tone flags or the red heart's variation selector) getting pulled out as separate emojis messes up your counts big time. Let's fix this properly.
The Problem with Your Original Code
Your current approach checks each character individually against emoji.UNICODE_EMOJI, but that set includes standalone modifier characters like Fitzpatrick skin tones (U+1F3FB to U+1F3FF) and variation selectors (like U+FE0F, the one in ❤️). That's why you're seeing split emojis in your output.
Solution 1: Use emoji.emoji_list() (Simplest & Most Reliable)
The emoji package has a built-in function designed to handle combined emoji units exactly how you need it. emoji_list() parses the text and returns full emoji entities—including those with modifiers—as single items, instead of splitting them into individual characters.
Here's how to rewrite your function:
import emoji def extract_emojis(text): # emoji_list returns a list of dicts; the 'emoji' key holds the full combined emoji return [item['emoji'] for item in emoji.emoji_list(text)]
Test It Out:
If you run:
print(extract_emojis("Hello ❤️👋🏿!"))
You'll get:
['❤️', '👋🏿']
Perfect—no split modifiers, just the complete emojis you care about.
Solution 2: Ignore Modifiers Entirely (Merge Skin Tone Variants)
If your goal is to treat emojis with different skin tones as the same base emoji (e.g., count 👋 and 👋🏿 as the same for stats), you can use emoji.demojize() to strip out the modifier tags, then convert back to the base emoji:
import emoji def extract_base_emojis(text): full_emojis = [item['emoji'] for item in emoji.emoji_list(text)] base_emojis = [] for e in full_emojis: # Convert emoji to text description (e.g., 👋🏿 becomes ":wave::skin-tone-5:") demojized = emoji.demojize(e) # Remove the skin tone modifier part base_desc = demojized.split(':skin-tone-')[0] + ':' # Convert back to the base emoji base_emoji = emoji.emojize(base_desc) base_emojis.append(base_emoji) return base_emojis
Test This Version:
Running:
print(extract_base_emojis("👋👋🏿❤️"))
Gives you:
['👋', '👋', '❤️']
Now you can count all skin tone variants of an emoji together.
Why Avoid Manual Code Point Filtering?
You could manually filter out the modifier code points (like U+1F3FB-U+1F3FF and U+FE0F), but that's error-prone—there are other less common modifiers you might miss, and you'd lose the context of combined emojis (e.g., you'd get 👋 instead of 👋🏿 if you filter the skin tone). Using the package's built-in functions is far more robust.
内容的提问来源于stack exchange,提问作者Alex

