You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python字符串分割问题:含特殊格式的课程信息拆分需求

解决课程信息字符串拆分问题

由于课程名称可能包含逗号,且存在多课程代码的情况,直接按, 拆分会导致字段错位,这里推荐使用正则表达式匹配固定格式的字段,精准提取所需内容。

Python实现方案

import re

def parse_course_string(course_str):
    # 定义正则匹配模式,使用VERBOSE模式提升可读性
    pattern = r'''
    ^(\d{4}-\d{2}-\d{2}),\s*                          # 匹配日期(YYYY-MM-DD格式)
    (\d{2}:\d{2})\s*-\s*(\d{2}:\d{2})\s*              # 匹配开始/结束时间(HH:MM - HH:MM)
    ([^,]+(?:,\s*[^,]+)*),\s*                         # 匹配课程代码,支持多代码用逗号+空格分隔
    (.*?)\s*,\s*(\w+)\s*,\s*                          # 匹配课程名称(允许含逗号)和课程类型(单个单词)
    ([^,]+),\s*                                       # 匹配教室名称
    (.*?)\s*,\s*([^,]+(?:,\s*[^,]+)*)                 # 匹配教学楼(允许含逗号)和分组(支持多分组)
    .*$                                               # 忽略末尾空字段
    '''
    match_result = re.match(pattern, course_str, re.VERBOSE)
    if not match_result:
        return None  # 匹配失败返回None
    
    # 提取字段并清理前后空格
    date = match_result.group(1).strip()
    start_time = match_result.group(2).strip()
    end_time = match_result.group(3).strip()
    course_code = match_result.group(4).strip()
    course_name = match_result.group(5).strip()
    course_type = match_result.group(6).strip()
    room = match_result.group(7).strip()
    building = match_result.group(8).strip()
    groups = match_result.group(9).strip()
    
    # 返回指定格式的列表
    return [date, start_time, end_time, course_code, course_name, course_type, room, building, groups]

# 测试示例
sample1 = "2024-03-21, 13:15 - 15:00 MVE545,  Mathematical analysis, part 2,  Lecture,  Omega,  0,  Jupiter,  TIDAL-1, TIELL-1,  ,  ,  ,  ,  ,"
sample2 = "2024-03-21, 13:15 - 15:00 MVE545, MVE545,  Mathematical analysis, part 2,  Lecture,  Omega,  0,  Jupiter,  TIDAL-1, TIELL-1,  ,  ,  ,  ,  ,"

print(parse_course_string(sample1))
# 输出: ['2024-03-21', '13:15', '15:00', 'MVE545', 'Mathematical analysis, part 2', 'Lecture', 'Omega', '0, Jupiter', 'TIDAL-1, TIELL-1']

print(parse_course_string(sample2))
# 输出: ['2024-03-21', '13:15', '15:00', 'MVE545, MVE545', 'Mathematical analysis, part 2', 'Lecture', 'Omega', '0, Jupiter', 'TIDAL-1, TIELL-1']

方案说明

  • 正则匹配逻辑:通过固定格式的正则规则,精准定位每个字段的边界,避免课程名称中的逗号干扰拆分。
  • 多场景支持:
    • 自动识别单个/多个课程代码(用逗号+空格分隔)
    • 正确提取含逗号的课程名称
    • 兼容末尾的空字段,不影响有效内容提取
  • 字段清理:通过strip()去除每个字段前后的多余空格,保证结果整洁。

内容的提问来源于stack exchange,提问作者Erik Carlsson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 15:45:25