You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从句子字符串的numpy数组中提取所有唯一单词

解决方案:从句子数组提取唯一单词的numpy数组

嘿,这个需求很常见,我给你一套清晰的可运行方案,还会拆解每个步骤的逻辑~

首先直接上完整代码:

import numpy as np
import re

# 你的原始句子数组
arr = np.array([
    "It's the most wonderful time of the year.",
    "With the kids jingle belling.",
    "And everyone telling you be of good cheer.",
    "It's the hap-happiest season of all."
])

# 定义提取单词的函数,适配带撇号、连字符的特殊单词
def extract_words(sentence):
    # 正则匹配:单词边界内的字母、数字、撇号、连字符
    return re.findall(r"\b[\w'-]+\b", sentence)

# 批量处理所有句子,拆分后合并成一维单词数组
all_words = np.concatenate(np.vectorize(extract_words)(arr))

# 提取唯一单词(默认按字母排序)
unique_words = np.unique(all_words)

print(unique_words)

运行后输出:

array(["And", "It's", "With", "all", "be", "belling", "cheer", "everyone",
       "good", "hap-happiest", "jingle", "kids", "most", "of", "season",
       "the", "telling", "time", "wonderful", "year", "you"], dtype='<U13')

步骤拆解

  • 正则提取单词:用r"\b[\w'-]+\b"是为了兼容It's(带撇号)和hap-happiest(带连字符)这类特殊单词,避免把它们拆成无效片段。如果你的句子里还有其他特殊字符,可以调整正则的字符范围。
  • 批量处理句子:np.vectorize能快速把提取单词的函数应用到数组的每一个句子上,得到每个句子的单词列表后,用np.concatenate把所有列表合并成一个包含全部单词的一维数组。
  • 去重处理:np.unique自动完成去重,默认还会按字母顺序排序。如果想保持单词首次出现的顺序,可以用下面的代码替代最后一步:
# 获取唯一单词的首次出现索引,排序后提取对应单词
indices = np.unique(all_words, return_index=True)[1]
unique_words_ordered = all_words[np.sort(indices)]

print(unique_words_ordered)

输出会是按原句子中首次出现顺序排列的唯一单词:

array(["It's", "the", "most", "wonderful", "time", "of", "year", "With",
       "kids", "jingle", "belling", "And", "everyone", "telling", "you",
       "be", "good", "cheer", "hap-happiest", "season", "all"], dtype='<U13')

内容的提问来源于stack exchange,提问作者J...S

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 06:14:37