如何在Python中提取字符串中的年份数字并添加到列表?
提取字符串中的年份数字到列表的Python实现
你的现有代码存在两个明显问题:
- 像
2000's这种带后缀的年份,会被split()当成完整单词,isdigit()无法识别,导致漏掉该年份; - 未限制数字位数,若字符串中出现其他长度的数字(比如"123"),会被误提取。
下面给出两种可行的解决方案:
方法一:正则表达式(最简洁精准)
正则能直接匹配所有4位数字的年份,不管前后是否带其他字符,用re.findall()可一次性提取所有符合条件的结果:
import re a = "Sam works in a company abc in New York. He joined the company last year 2019. Before joining ABC, he used to work for a small firm in Arizona. He worked there from 2015 to 2018. Before moving to Arizona Sam used to live in South Dakota and he has been living there since 2000's" # 匹配所有4位数字,返回字符串列表 years = re.findall(r'\b\d{4}\b', a) # 如需转成整数列表 years_int = [int(year) for year in years] print(years) # 输出: ['2019', '2015', '2018', '2000'] print(years_int) # 输出: [2019, 2015, 2018, 2000]
正则r'\b\d{4}\b'的说明:
\b是单词边界,确保匹配的是完整的4位数字(不会匹配"12345"里的前4位);\d{4}表示匹配连续4个数字;re.findall()返回所有符合条件的匹配项组成的列表。
方法二:改进原始循环方法
如果不想用正则,可以修改原代码,先清理单词中的非数字字符,再判断是否为4位数字:
a = "Sam works in a company abc in New York. He joined the company last year 2019. Before joining ABC, he used to work for a small firm in Arizona. He worked there from 2015 to 2018. Before moving to Arizona Sam used to live in South Dakota and he has been living there since 2000's" b = a.split() year = [] for i in b: # 提取单词中的所有数字字符 digits = ''.join([c for c in i if c.isdigit()]) # 判断是否是4位数字 if len(digits) == 4 and digits.isdigit(): year.append(digits) print(year) # 输出: ['2019', '2015', '2018', '2000']
该方法先抽取每个单词里的数字,再检查长度是否为4,能处理2000's这类带后缀的情况。
内容的提问来源于stack exchange,提问作者DevWms27
相关产品推荐
相关产品推荐

