Python字符串分割问题:含特殊格式的课程信息拆分需求
解决课程信息字符串拆分问题
由于课程名称可能包含逗号,且存在多课程代码的情况,直接按, 拆分会导致字段错位,这里推荐使用正则表达式匹配固定格式的字段,精准提取所需内容。
Python实现方案
import re def parse_course_string(course_str): # 定义正则匹配模式,使用VERBOSE模式提升可读性 pattern = r''' ^(\d{4}-\d{2}-\d{2}),\s* # 匹配日期(YYYY-MM-DD格式) (\d{2}:\d{2})\s*-\s*(\d{2}:\d{2})\s* # 匹配开始/结束时间(HH:MM - HH:MM) ([^,]+(?:,\s*[^,]+)*),\s* # 匹配课程代码,支持多代码用逗号+空格分隔 (.*?)\s*,\s*(\w+)\s*,\s* # 匹配课程名称(允许含逗号)和课程类型(单个单词) ([^,]+),\s* # 匹配教室名称 (.*?)\s*,\s*([^,]+(?:,\s*[^,]+)*) # 匹配教学楼(允许含逗号)和分组(支持多分组) .*$ # 忽略末尾空字段 ''' match_result = re.match(pattern, course_str, re.VERBOSE) if not match_result: return None # 匹配失败返回None # 提取字段并清理前后空格 date = match_result.group(1).strip() start_time = match_result.group(2).strip() end_time = match_result.group(3).strip() course_code = match_result.group(4).strip() course_name = match_result.group(5).strip() course_type = match_result.group(6).strip() room = match_result.group(7).strip() building = match_result.group(8).strip() groups = match_result.group(9).strip() # 返回指定格式的列表 return [date, start_time, end_time, course_code, course_name, course_type, room, building, groups] # 测试示例 sample1 = "2024-03-21, 13:15 - 15:00 MVE545, Mathematical analysis, part 2, Lecture, Omega, 0, Jupiter, TIDAL-1, TIELL-1, , , , , ," sample2 = "2024-03-21, 13:15 - 15:00 MVE545, MVE545, Mathematical analysis, part 2, Lecture, Omega, 0, Jupiter, TIDAL-1, TIELL-1, , , , , ," print(parse_course_string(sample1)) # 输出: ['2024-03-21', '13:15', '15:00', 'MVE545', 'Mathematical analysis, part 2', 'Lecture', 'Omega', '0, Jupiter', 'TIDAL-1, TIELL-1'] print(parse_course_string(sample2)) # 输出: ['2024-03-21', '13:15', '15:00', 'MVE545, MVE545', 'Mathematical analysis, part 2', 'Lecture', 'Omega', '0, Jupiter', 'TIDAL-1, TIELL-1']
方案说明
- 正则匹配逻辑:通过固定格式的正则规则,精准定位每个字段的边界,避免课程名称中的逗号干扰拆分。
- 多场景支持:
- 自动识别单个/多个课程代码(用逗号+空格分隔)
- 正确提取含逗号的课程名称
- 兼容末尾的空字段,不影响有效内容提取
- 字段清理:通过
strip()去除每个字段前后的多余空格,保证结果整洁。
内容的提问来源于stack exchange,提问作者Erik Carlsson
相关产品推荐
相关产品推荐

