You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何对归一化后的短语执行二次分词生成对应列?

问题:归一化短语的二次分词实现

现有数据集预处理流程为:初次分词→俚语归一化,但归一化后的俚语可能是带空格的短语,需要实现二次分词生成secondTokenization列。数据示例如下:

firstTokenization                       normalized               secondTokenization 
0     [yes, no, cs]      [yes, no, customer service]     [yes, no, customer, service] 
1             [nlp]    [natural language processing]  [natural, language, processing] 
2         [no, yes]                        [no, yes]                        [no, yes] 

当前已完成初次分词和归一化的代码如下:

tokenizer = MWETokenizer()
def tokenization (text):
    return tokenizer.tokenize(text.split())
df['firstTokenization'] = df['content'].apply(lambda x: tokenization(x.lower()))

normalizad_word = pd.read_excel('normalisasi.xlsx')
normalizad_word_dict = {}

for index, row in normalizad_word.iterrows():
    if row[0] not in normalizad_word_dict:
        normalizad_word_dict[row[0]] = row[1] 
def normalized_term(document):
    return [normalizad_word_dict[term] if term in normalizad_word_dict else term for term in document]
df['normalized'] = df['firstTokenization'].apply(normalized_term)

解决方案

只需新增一个二次分词函数,遍历normalized列的每个元素,将带空格的短语拆分为单个词,再合并为扁平列表即可:

def secondary_tokenization(normalized_doc):
    final_tokens = []
    for term in normalized_doc:
        # 按空格拆分每个归一化项,扩展到结果列表中
        final_tokens.extend(term.split())
    return final_tokens

# 生成secondTokenization列
df['secondTokenization'] = df['normalized'].apply(secondary_tokenization)

说明

  • term.split()会自动将带空格的短语拆分为独立词汇,比如"customer service"会拆成["customer", "service"]
  • extend()方法会把拆分后的词汇逐个加入结果列表,避免生成嵌套列表,最终得到扁平的分词结果,完全匹配需求中的secondTokenization格式。

内容的提问来源于stack exchange,提问作者Dewani

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 08:55:21