You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python将无空格英文数字表述字符串转换为整数?

无空格数字英文串转整数的起步实现方案

先修正基础字典

你提供的num2words字典存在两个问题:一是值的类型混合(有字符串也有整数),二是存在拼写错误(fourty应为forty),先统一修正为全整数类型的正确字典:

num2words = {
    # 个位数
    'one': 1, 'two': 2, 'three': 3, 'four':4, 'five':5, 'six':6, 'seven':7, 'eight':8, 'nine':9,
    # 10-19特殊数
    'ten':10, 'eleven':11, 'twelve':12, 'thirteen':13, 'fourteen':14, 'fifteen':15, 'sixteen':16,
    'seventeen':17, 'eighteen':18, 'nineteen':19,
    # 整十数
    'twenty':20, 'thirty':30, 'forty':40, 'fifty':50, 'sixty':60, 'seventy':70, 'eighty':80,
    'ninety':90,
    # 量级词
    'hundred':100, 'thousand':1000, 'million':1000000, 'zero':0
}

核心步骤1:拆分无空格字符串

无空格串的核心问题是把连续字符拆成合法的数字单词,这里用贪心匹配最长单词的思路:优先尝试匹配字典里最长的单词,避免短单词误拆分(比如twentyseven不会被拆成twen+tyseven,而是优先匹配twenty再匹配seven)。

实现拆分函数:

def split_num_string(s, word_dict):
    # 按单词长度从长到短排序,优先匹配长单词
    sorted_words = sorted(word_dict.keys(), key=lambda x: -len(x))
    result = []
    current_pos = 0
    str_len = len(s)
    
    while current_pos < str_len:
        matched = False
        for word in sorted_words:
            word_len = len(word)
            # 检查当前位置起的子串是否匹配单词
            if current_pos + word_len <= str_len and s[current_pos:current_pos+word_len] == word:
                result.append(word)
                current_pos += word_len
                matched = True
                break
        if not matched:
            raise ValueError(f"无法识别的子串:{s[current_pos:]}")
    return result

核心步骤2:拆分后的单词转整数

这部分逻辑和带空格的场景一致,需要处理hundred、thousand这类量级词的乘法规则:

  • 用current_total保存当前未遇到量级词的累加值
  • 遇到hundred时,将current_total乘以100
  • 遇到thousand时,将current_total乘以1000后加到总结果,再重置current_total
  • 最后把剩余的current_total加到总结果

实现转换函数:

def words_to_number(word_list, word_dict):
    current_total = 0
    grand_total = 0
    
    for word in word_list:
        val = word_dict[word]
        if val == 100:
            current_total *= val
        elif val == 1000:
            grand_total += current_total * val
            current_total = 0
        else:
            current_total += val
    # 加上最后一段未处理的数值
    grand_total += current_total
    return grand_total

整合并测试

把拆分和转换逻辑整合,写一个主函数并测试:

def num_str_to_int(s):
    # 初始化修正后的字典
    num2words = {
        'one': 1, 'two': 2, 'three': 3, 'four':4, 'five':5, 'six':6, 'seven':7, 'eight':8, 'nine':9,
        'ten':10, 'eleven':11, 'twelve':12, 'thirteen':13, 'fourteen':14, 'fifteen':15, 'sixteen':16,
        'seventeen':17, 'eighteen':18, 'nineteen':19,
        'twenty':20, 'thirty':30, 'forty':40, 'fifty':50, 'sixty':60, 'seventy':70, 'eighty':80,
        'ninety':90,
        'hundred':100, 'thousand':1000, 'million':1000000, 'zero':0
    }
    # 拆分字符串
    word_list = split_num_string(s, num2words)
    # 转换为整数
    return words_to_number(word_list, num2words)

# 测试用例
print(num_str_to_int("twentyseven"))  # 输出27
print(num_str_to_int("threehundredfortyfive"))  # 输出345
print(num_str_to_int("onethousandtwohundredthirtyfour"))  # 输出1234

额外注意事项

  • 若输入可能包含连字符(比如twenty-seven),可先做预处理:s = s.replace('-', '')
  • 可以增加输入合法性校验(比如检查输入是否全小写、是否存在无法匹配的单词)
  • 针对小于1,000,000的限制,可以在最后判断结果是否符合范围,超出则抛出异常

内容的提问来源于stack exchange,提问作者KenFuzion

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 18:21:27