如何对归一化后的短语执行二次分词生成对应列?
问题:归一化短语的二次分词实现
现有数据集预处理流程为:初次分词→俚语归一化,但归一化后的俚语可能是带空格的短语,需要实现二次分词生成secondTokenization列。数据示例如下:
firstTokenization normalized secondTokenization 0 [yes, no, cs] [yes, no, customer service] [yes, no, customer, service] 1 [nlp] [natural language processing] [natural, language, processing] 2 [no, yes] [no, yes] [no, yes]
当前已完成初次分词和归一化的代码如下:
tokenizer = MWETokenizer() def tokenization (text): return tokenizer.tokenize(text.split()) df['firstTokenization'] = df['content'].apply(lambda x: tokenization(x.lower())) normalizad_word = pd.read_excel('normalisasi.xlsx') normalizad_word_dict = {} for index, row in normalizad_word.iterrows(): if row[0] not in normalizad_word_dict: normalizad_word_dict[row[0]] = row[1] def normalized_term(document): return [normalizad_word_dict[term] if term in normalizad_word_dict else term for term in document] df['normalized'] = df['firstTokenization'].apply(normalized_term)
解决方案
只需新增一个二次分词函数,遍历normalized列的每个元素,将带空格的短语拆分为单个词,再合并为扁平列表即可:
def secondary_tokenization(normalized_doc): final_tokens = [] for term in normalized_doc: # 按空格拆分每个归一化项,扩展到结果列表中 final_tokens.extend(term.split()) return final_tokens # 生成secondTokenization列 df['secondTokenization'] = df['normalized'].apply(secondary_tokenization)
说明
term.split()会自动将带空格的短语拆分为独立词汇,比如"customer service"会拆成["customer", "service"]extend()方法会把拆分后的词汇逐个加入结果列表,避免生成嵌套列表,最终得到扁平的分词结果,完全匹配需求中的secondTokenization格式。
内容的提问来源于stack exchange,提问作者Dewani
相关产品推荐
相关产品推荐

