基于NLTK的NLP计时器多时间单位组合识别开发咨询
基于Python与NLTK的多单位自然语言计时器实现方案
原代码无法识别组合时长的核心原因是,小时、分钟、秒的判断逻辑完全独立,每次匹配到单位就把句子里所有剩余数字全塞给该单位,没有建立数字和相邻单位的对应关系,只要句子里同时出现两个时间单位就会出现数值匹配错误。
改造核心逻辑
顺着分词后的文本顺序逐词扫描即可,逻辑非常适合初学者理解:
核心匹配规则:数字永远绑定紧跟在它后面的第一个时间单位,符合日常自然语言表达习惯,不会出现匹配错位。
- 提前定义好所有需要过滤的无意义停用词,以及时间单位的匹配规则
- 初始化小时、分钟、秒的默认值为0,再加一个临时变量暂存当前扫描到的数字
- 逐词遍历处理好的文本:
- 碰到数字就先存到临时变量里
- 碰到时间单位,就把之前暂存的数字赋值给对应单位,赋值完清空临时变量
- 碰到停用词直接跳过
- 最后把匹配到的三个时间值传给原有的倒计时函数即可,不需要改动倒计时逻辑
改造后完整代码
import time import datetime from nltk_utils import bag_of_words, tokenize import multiprocessing from playsound import playsound def is_number(x): if type(x) == str: x = x.replace(',', '') try: float(x) except: return False return True def text2int (textnum, numwords={}): units = [ 'zero', 'one', 'two', 'three', 'four', 'five', 'six', 'seven', 'eight', 'nine', 'ten', 'eleven', 'twelve', 'thirteen', 'fourteen', 'fifteen', 'sixteen', 'seventeen', 'eighteen', 'nineteen', ] tens = ['', '', 'twenty', 'thirty', 'forty', 'fifty', 'sixty', 'seventy', 'eighty', 'ninety'] scales = ['hundred', 'thousand', 'million', 'billion', 'trillion'] ordinal_words = {'first':1, 'second':2, 'third':3, 'fifth':5, 'eighth':8, 'ninth':9, 'twelfth':12} ordinal_endings = [('ieth', 'y'), ('th', '')] if not numwords: numwords['and'] = (1, 0) for idx, word in enumerate(units): numwords[word] = (1, idx) for idx, word in enumerate(tens): numwords[word] = (1, idx * 10) for idx, word in enumerate(scales): numwords[word] = (10 ** (idx * 3 or 2), 0) textnum = textnum.replace('-', ' ') current = result = 0 curstring = '' onnumber = False lastunit = False lastscale = False def is_numword(x): if is_number(x): return True if word in numwords: return True return False def from_numword(x): if is_number(x): scale = 0 increment = int(x.replace(',', '')) return scale, increment return numwords[x] for word in textnum.split(): if word in ordinal_words: scale, increment = (1, ordinal_words[word]) current = current * scale + increment if scale > 100: result += current current = 0 onnumber = True lastunit = False lastscale = False else: for ending, replacement in ordinal_endings: if word.endswith(ending): word = "%s%s" % (word[:-len(ending)], replacement) if (not is_numword(word)) or (word == 'and' and not lastscale): if onnumber: curstring += repr(result + current) + " " curstring += word + " " result = current = 0 onnumber = False lastunit = False lastscale = False else: scale, increment = from_numword(word) onnumber = True if lastunit and (word not in scales): curstring += repr(result + current) result = current = 0 if scale > 1: current = max(1, current) current = current * scale + increment if scale > 100: result += current current = 0 lastscale = False lastunit = False if word in scales: lastscale = True elif word in units: lastunit = True if onnumber: curstring += repr(result + current) return curstring # 统一抽离停用词表,避免重复编写相同内容 STOP_WORDS = {'second', 'seconds', 'minute', 'minutes', 'hour', 'hours', 'please', 'set', 's', 'timer', '``', '{', '}', 'text', ':', "''", 'a', 'for', 'and', 'if', 'privacy', 'time', 'but', 'end', 'put', 'me', 'my', 'will', 'you', 'now', 'right', 'privacy', 'rite', 'wright', 'write', 'your', 'go', 'ahead', 't'} def countdown(h, m, s): total_seconds = h * 3600 + m * 60 + s while total_seconds > 0: timer = datetime.timedelta(seconds = total_seconds) print(timer, end="\r") time.sleep(1) total_seconds -= 1 print('timer ended') if __name__ == "__main__": # 测试输入,支持多单位组合、英文数词、阿拉伯数字混合输入 input_text = "please will you set a 10 minutes 45 seconds timer" sentence_to_num = text2int(input_text) tokens = tokenize(sentence_to_num) print("分词结果:", tokens) # 初始化时间值和临时数字存储 h = m = s = 0 current_num = None for word in tokens: # 碰到数字先暂存 if is_number(word): current_num = int(float(word.replace(',', ''))) continue # 碰到时间单位就绑定前面暂存的数字 if word in ["hour", "hours"] and current_num is not None: h = current_num current_num = None print(f"{h} hours starting now") elif word in ["minute", "minutes"] and current_num is not None: m = current_num current_num = None print(f"{m} minutes starting now") elif word in ["second", "seconds"] and current_num is not None: s = current_num current_num = None print(f"{s} seconds starting now") # 其他停用词直接跳过 else: continue # 启动倒计时 countdown(h, m, s)
效果说明
- 支持纯数字输入(如
10 minutes 45 seconds)、英文数词输入(如ten minutes forty five seconds)、多单位混合输入(如1 hour 5 minutes 30 seconds) - 自动跳过所有无意义的停用词,不会出现原代码里数字重复赋值给多个单位的问题
- 代码逻辑线性简单,后续要加天、毫秒等其他时间单位,只需要在单位判断的分支里加对应规则即可
内容的提问来源于stack exchange,提问作者Brandon Jones
相关产品推荐
相关产品推荐

