如何在Python Pandas中实现Emoji转文本(字典值替换为键)
Hey there! Let's work through this emoji-to-text replacement task in Pandas with an efficient, clean approach. Your dictionary has some multi-emoji entries, so we'll first get that sorted into a usable mapping, then use a regex-powered replacement that's way faster than looping through each emoji one by one.
Step 1: Fix and Transform Your Dictionary
First, let's clean up your original dictionary (I fixed the quote syntax to make it valid Python) and convert it into an emoji-to-text mapping—since we need to look up emojis and replace them with their corresponding labels. This handles cases where multiple emojis map to the same text:
import pandas as pd import re # Your original dictionary (fixed for valid syntax) original_emoji_dict = { 'butterfly': "Ƹ̵̡Ӝ̵̨̄Ʒ", 'clapping hands': "o/*, *o/*", 'face with raised eyebrow': "O?O", 'face with symbols on mouth': ">.'", 'grimacing face': "e.e, O.e, O.e", 'rolling on the floor laughing': "m/*.*m/" } # Build a reversed emoji-to-text mapping emoji_to_text = {} for label, emoji_str in original_emoji_dict.items(): # Split comma-separated emojis, strip extra whitespace/quotes emojis = [emoji.strip().strip("'") for emoji in emoji_str.split(',')] for emoji in emojis: emoji_to_text[emoji] = label
Step 2: Efficient Replacement in Pandas
Instead of running multiple str.replace calls (which gets slow with big datasets), we'll create a single regex pattern that matches all emojis in our mapping. Pandas can use this pattern with a lambda function to replace every match in one go:
# Your sample data sample_data = pd.DataFrame({ 'text': [ ".@AnnaKendrick47 My set up at the electronics boat at work. ^_^ \"Fun update for everyone who's requested, #EW is now IN!! @WordsWFriends ⬇️ ⬇️ ⬇️ '@AnnaKendrick47 please sing @DrewGasparini 's Circus\"\"'" ] }) # Create a regex pattern that matches any emoji in our mapping # re.escape ensures special characters (like *, /) don't break the regex emoji_pattern = re.compile('|'.join(re.escape(emoji) for emoji in emoji_to_text.keys())) # Replace emojis with their labels sample_data['clean_text'] = sample_data['text'].str.replace( emoji_pattern, lambda match: emoji_to_text[match.group()], regex=True ) # Check the result print(sample_data['clean_text'].iloc[0])
Why This Approach Rocks
- Speed: A single regex scan is way more efficient than looping through each emoji and replacing individually—critical if you're working with large datasets.
- Handles Multi-Emoji Labels: Properly maps every emoji (even duplicates like the grimacing face entries) to the correct text label.
- Robust: Uses
re.escapeto handle special characters in emojis (like*or?) so your regex doesn't break unexpectedly.
If your sample data included any emojis from your dictionary (say, Ƹ̵̡Ӝ̵̨̄Ʒ), it would get replaced with butterfly automatically. The downward arrows (⬇️) aren't in your mapping, so they stay as-is.
内容的提问来源于stack exchange,提问作者raj chejara

