如何从字符串中提取数字的字母表示(Python 3.x)
从文本中提取单词形式的数字(含复合数字)
这需求挺实用的——要从自然语言文本里揪出那些用单词表示的数字,不管是单个的(比如five)还是多词组合的(比如sixty six),而不是把它们转成数值。下面我给你一个Python的实现方案,完全符合你要的find_numbers函数效果:
实现思路
- 定义完整的数字单词库:覆盖0-19的基础数字词、几十(twenty到ninety),如果需要还能扩展更高位的词(比如hundred、thousand);
- 用正则匹配连续数字词:利用正则的单词边界和重复匹配规则,捕获连续出现的数字单词组合,这样就能拿到像
sixty six这类复合数字; - 封装成函数:用正则的
findall方法提取所有符合条件的子串,再整理成清晰的结果列表。
代码实现
import re def find_numbers(text): # 定义所有数字相关的单词,可按需扩展(比如加上hundred、thousand等) number_words = [ 'zero', 'one', 'two', 'three', 'four', 'five', 'six', 'seven', 'eight', 'nine', 'ten', 'eleven', 'twelve', 'thirteen', 'fourteen', 'fifteen', 'sixteen', 'seventeen', 'eighteen', 'nineteen', 'twenty', 'thirty', 'forty', 'fifty', 'sixty', 'seventy', 'eighty', 'ninety' ] # 构建正则模式:匹配一个或多个连续的数字单词,单词间用空格分隔 pattern = r'\b(' + '|'.join(number_words) + r')(?:\s+(' + '|'.join(number_words) + r'))*\b' # 提取所有匹配结果,处理分组拼接成完整数字短语 matches = re.findall(pattern, text.lower()) result = [] for match in matches: full_num = ' '.join(filter(None, match)) if full_num: result.append(full_num) return result
测试验证
跑几个例子看看效果:
print(find_numbers("Hello world")) # 输出: [] print(find_numbers("Hello five world")) # 输出: ['five'] print(find_numbers("I've got sixty six tasks")) # 输出: ['sixty six'] print(find_numbers("There is four people and twenty five cats")) # 输出: ['four', 'twenty five'] print(find_numbers("One hundred thirty five apples")) # 输出: ['one hundred thirty five'](需在number_words中添加hundred)
注意事项
- 若要处理更大的数字(比如hundred、thousand),只需把对应单词加入
number_words列表即可; - 正则里的
\b是单词边界,能避免误匹配包含数字单词的普通词汇(比如不会把fivetree里的five揪出来); - 函数里转小写是为了不区分大小写(比如
Five和five都能被匹配),如果需要严格区分大小写,去掉.lower()即可。
内容的提问来源于stack exchange,提问作者Богдан Марченко
相关产品推荐
相关产品推荐

