Haskell实现移除高频英文词功能:如何获取常用词列表?
实现dropCommonWords函数的思路与方案
没错,你确实需要自行创建英文前20个最常用词的列表——Haskell的标准库并没有内置这个集合,得咱们自己定义。
首先,先把基于通用语料统计的经典前20个高频常用词列出来:
commonWords :: [String] commonWords = ["the", "be", "to", "of", "and", "a", "in", "that", "have", "I", "it", "for", "not", "on", "with", "he", "as", "you", "do", "at"]
接下来,参考你写的dropletters的风格,咱们可以用filter函数实现dropCommonWords,核心逻辑就是保留不在常用词列表里的元素:
dropCommonWords :: [String] -> [String] dropCommonWords xs = filter (\x -> x `notElem` commonWords) xs
测试你给的示例:
dropCommonWords ["the","planet","of","the","apes"]
["planet","apes"]
完全符合预期~
如果需要支持不区分大小写的过滤(比如输入"The"也能被移除),可以借助Data.Char里的toLower函数,把所有词统一转成小写再判断:
import Data.Char (toLower) dropCommonWordsIgnoreCase :: [String] -> [String] dropCommonWordsIgnoreCase xs = filter (\x -> map toLower x `notElem` lowerCommonWords) xs where lowerCommonWords = map (map toLower) commonWords
这样不管输入是大写、小写还是混合大小写,都能正确过滤掉常用词啦。
内容的提问来源于stack exchange,提问作者lylyly
相关产品推荐
相关产品推荐

