You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中如何基于表情符号拆分解析含字符串、数字的评论内容

实现方案

核心通过完整匹配全量表情符号作为分隔符的方式拆分字符串,不会残留表情组成字符,通用实现代码如下:

import re

# 先定义所有需要识别的表情正则规则,注意长表情放前面避免短表情优先匹配错误
EMOJI_PATTERN = re.compile(r'(?:>:O|:\)|:O|:v)')

def split_comment(comment: str) -> list[str]:
    # 按表情拆分字符串,对每个片段去除首尾空格,过滤空片段
    parts = [part.strip() for part in EMOJI_PATTERN.split(comment)]
    return [p for p in parts if p]

# 测试示例
comment_1 = "This is :) my comment :O"
comment_2 = ">:O Another comment to :v parse"

output_1 = split_comment(comment_1)
output_2 = split_comment(comment_2)

print(output_1) # 输出:['This is', 'my comment']
print(output_2) # 输出:['Another comment to', 'parse']

扩展说明

  • 正则规则用非捕获组(?:)包裹所有表情,拆分时会把完整表情整体作为分隔符移除,不会残留O/v这类表情组成字符
  • 新增表情时直接把表情字符串(正则特殊字符加\转义)加入正则的|分隔列表即可,注意长度更长的表情要放在更靠前的位置,避免短表情提前匹配截断长表情

内容的提问来源于stack exchange,提问作者The Dan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.07 07:15:03