如何编写正则表达式匹配指定年份范围的系列ID?
匹配201x/202x开头的系列ID并提取年份、保存结果
正则匹配规则
- 筛选符合条件的ID:使用
^20[12]\d,确保ID前四位是2010-2029范围内的年份开头 - 提取年份:用
^(\d{4})直接截取ID的前四位作为年份
完整代码实现
import re def process_ids(input_file, output_file): id_pattern = r'^20[12]\d' year_pattern = r'^(\d{4})' with open(input_file, 'r') as infile, open(output_file, 'w') as outfile: for line in infile: line = line.strip() if re.match(id_pattern, line): year = re.findall(year_pattern, line)[0] # 按需选择写入格式,以下是两种示例: # 仅保存ID:outfile.write(f"{line}\n") # 保存ID+年份: outfile.write(f"ID: {line}, 年份: {year}\n") # 调用示例,替换为你的文件路径 process_ids("source.txt", "filtered_result.txt")
代码说明
^20[12]\d:^锚定字符串开头,20[12]匹配"201"或"202",\d匹配任意数字,精准筛选目标ID- 逐行读取原始文件,跳过空行,符合条件的行提取年份后写入新文件
- 可根据需求修改写入格式,比如只保留ID或单独存储年份
样本数据处理结果
样本中符合条件的ID如下:
2012030116
2012030395
2012030396
2012030364
2012030246
提取的年份均为2012,处理后会写入指定的输出文件。
内容的提问来源于stack exchange,提问作者Store LTE
相关产品推荐
相关产品推荐

