You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python Pandas中实现Emoji转文本(字典值替换为键)

Hey there! Let's work through this emoji-to-text replacement task in Pandas with an efficient, clean approach. Your dictionary has some multi-emoji entries, so we'll first get that sorted into a usable mapping, then use a regex-powered replacement that's way faster than looping through each emoji one by one.

Step 1: Fix and Transform Your Dictionary

First, let's clean up your original dictionary (I fixed the quote syntax to make it valid Python) and convert it into an emoji-to-text mapping—since we need to look up emojis and replace them with their corresponding labels. This handles cases where multiple emojis map to the same text:

import pandas as pd
import re

# Your original dictionary (fixed for valid syntax)
original_emoji_dict = {
    'butterfly': "Ƹ̵̡Ӝ̵̨̄Ʒ",
    'clapping hands': "o/*, *o/*",
    'face with raised eyebrow': "O?O",
    'face with symbols on mouth': ">.'",
    'grimacing face': "e.e, O.e, O.e",
    'rolling on the floor laughing': "m/*.*m/"
}

# Build a reversed emoji-to-text mapping
emoji_to_text = {}
for label, emoji_str in original_emoji_dict.items():
    # Split comma-separated emojis, strip extra whitespace/quotes
    emojis = [emoji.strip().strip("'") for emoji in emoji_str.split(',')]
    for emoji in emojis:
        emoji_to_text[emoji] = label

Step 2: Efficient Replacement in Pandas

Instead of running multiple str.replace calls (which gets slow with big datasets), we'll create a single regex pattern that matches all emojis in our mapping. Pandas can use this pattern with a lambda function to replace every match in one go:

# Your sample data
sample_data = pd.DataFrame({
    'text': [
        ".@AnnaKendrick47 My set up at the electronics boat at work. ^_^ \"Fun update for everyone who's requested, #EW is now IN!! @WordsWFriends ⬇️ ⬇️ ⬇️ '@AnnaKendrick47 please sing @DrewGasparini 's Circus\"\"'"
    ]
})

# Create a regex pattern that matches any emoji in our mapping
# re.escape ensures special characters (like *, /) don't break the regex
emoji_pattern = re.compile('|'.join(re.escape(emoji) for emoji in emoji_to_text.keys()))

# Replace emojis with their labels
sample_data['clean_text'] = sample_data['text'].str.replace(
    emoji_pattern,
    lambda match: emoji_to_text[match.group()],
    regex=True
)

# Check the result
print(sample_data['clean_text'].iloc[0])

Why This Approach Rocks

  • Speed: A single regex scan is way more efficient than looping through each emoji and replacing individually—critical if you're working with large datasets.
  • Handles Multi-Emoji Labels: Properly maps every emoji (even duplicates like the grimacing face entries) to the correct text label.
  • Robust: Uses re.escape to handle special characters in emojis (like * or ?) so your regex doesn't break unexpectedly.

If your sample data included any emojis from your dictionary (say, Ƹ̵̡Ӝ̵̨̄Ʒ), it would get replaced with butterfly automatically. The downward arrows (⬇️) aren't in your mapping, so they stay as-is.

内容的提问来源于stack exchange,提问作者raj chejara

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:27:09