如何从句子字符串的numpy数组中提取所有唯一单词
解决方案:从句子数组提取唯一单词的numpy数组
嘿,这个需求很常见,我给你一套清晰的可运行方案,还会拆解每个步骤的逻辑~
首先直接上完整代码:
import numpy as np import re # 你的原始句子数组 arr = np.array([ "It's the most wonderful time of the year.", "With the kids jingle belling.", "And everyone telling you be of good cheer.", "It's the hap-happiest season of all." ]) # 定义提取单词的函数,适配带撇号、连字符的特殊单词 def extract_words(sentence): # 正则匹配:单词边界内的字母、数字、撇号、连字符 return re.findall(r"\b[\w'-]+\b", sentence) # 批量处理所有句子,拆分后合并成一维单词数组 all_words = np.concatenate(np.vectorize(extract_words)(arr)) # 提取唯一单词(默认按字母排序) unique_words = np.unique(all_words) print(unique_words)
运行后输出:
array(["And", "It's", "With", "all", "be", "belling", "cheer", "everyone", "good", "hap-happiest", "jingle", "kids", "most", "of", "season", "the", "telling", "time", "wonderful", "year", "you"], dtype='<U13')
步骤拆解
- 正则提取单词:用
r"\b[\w'-]+\b"是为了兼容It's(带撇号)和hap-happiest(带连字符)这类特殊单词,避免把它们拆成无效片段。如果你的句子里还有其他特殊字符,可以调整正则的字符范围。 - 批量处理句子:
np.vectorize能快速把提取单词的函数应用到数组的每一个句子上,得到每个句子的单词列表后,用np.concatenate把所有列表合并成一个包含全部单词的一维数组。 - 去重处理:
np.unique自动完成去重,默认还会按字母顺序排序。如果想保持单词首次出现的顺序,可以用下面的代码替代最后一步:
# 获取唯一单词的首次出现索引,排序后提取对应单词 indices = np.unique(all_words, return_index=True)[1] unique_words_ordered = all_words[np.sort(indices)] print(unique_words_ordered)
输出会是按原句子中首次出现顺序排列的唯一单词:
array(["It's", "the", "most", "wonderful", "time", "of", "year", "With", "kids", "jingle", "belling", "And", "everyone", "telling", "you", "be", "good", "cheer", "hap-happiest", "season", "all"], dtype='<U13')
内容的提问来源于stack exchange,提问作者J...S
相关产品推荐
相关产品推荐

