Spark使用flatMap调用pos_tag_counter处理文本词性标注结果异常
问题分析
- 出现单个单词词性标注结果的核心原因是
pos_tag_counter函数当前的输出不符合预期:该函数需要输入一行文本后,返回整行识别完成的词组+词性元组列表,格式为[('events of the day', 'NP'), ('last season', 'NN')...],如果当前返回的是单个单词的词性元组,说明函数内部没有实现词组组块识别逻辑,需要先修正该函数。 - 现有统计词性频次的逻辑本身是通顺的,只要
phrasesRDD的元素为(词组, 词性)格式,后续的映射和聚合操作就能正常统计词性出现次数。
修正后完整代码
import re # 正则匹配过滤URL开头行 reg = re.compile('^(?!URL).*') # 步骤1:过滤空行和URL开头行,增加strip()处理排除全空格的无效行 non_empty_lines = text.filter(lambda x: len(x.strip()) > 0) no_urls = non_empty_lines.filter(lambda x: reg.match(x.strip())) # 步骤2:按行处理生成词组+词性元组 # 注意:必须保证pos_tag_counter返回[(词组1, 词性1), (词组2, 词性2)...]格式 phrases = no_urls.flatMap(lambda line: pos_tag_counter(line)) # 步骤3:统计词性频次 part_of_speech = phrases.map(lambda item: (item[1], 1)) pos_count = part_of_speech.reduceByKey(lambda accumulator, value: accumulator + value) # 步骤4:按频次降序排序,得到最终结果 sorted_pos = pos_count.sortBy(lambda x: x[1], ascending=False) # 输出结果 print(sorted_pos.collect())
附加说明
如果你的pos_tag_counter是直接调用nltk等工具库的原生pos_tag接口,默认返回的就是单个单词的词性标注结果,要实现词组级的词性标注,需要额外补充名词短语组块(NP Chunking)逻辑,比如基于正则表达式匹配词性序列实现组块提取。
内容的提问来源于stack exchange,提问作者Hefe
相关产品推荐
相关产品推荐

