You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中优雅实现多方法尝试直至成功的优化方案问询(HTML日期提取)

优雅提取HTML中多格式日期的Python实现

我需要从HTML页面提取最后修改日期这类信息,但该信息存在多种非统一的声明方式,无法通过简单循环处理。目前有两种不太理想的实现方式:

第一种是逐次判断的繁琐写法:

def get_date(html):
  date = None
  # Approach 1
  time_tag = html.find("time", {"datetime": True})
  if time_tag:
    date = time_tag["datetime"]
  if date:
    return date

  # Approach 2
  mod_tag = html.find("meta", {"property": "article:modified_time", "content": True})
  if mod_tag:
    date = mod_tag["content"]
  if date:
    return date
  
  # Approach 3
  div_tag = html.find("div", {"class": "dateline"})
  if div_tag:
    date = div_tag.get_text()
  if date:
    return date
  
  # Approach n
  # ...

  return date

第二种是将每种提取方法封装为函数后循环调用,但这些函数仅单次使用,会增加代码量且污染命名空间:

def method_1(html):
  test = html.find("time", {"datetime": True})
  return test["datetime"] if test else None

def method_2(html):
  test = html.find("meta", {"property": "article:modified_time", "content": True})
  return test["content"] if test else None

def method_3(html):
  test = html.find("div", {"class": "dateline"})
  return test.get_text() if test else None

...

def get_date(html):
  date = None
  bag_of_methods = [method_1, method_2, method_3, ...]

  i = 0
  while not date and i < len(bag_of_methods):
    date = bag_of_methods[i](html)
    i += 1

  return date

请问在Python中是否存在更简洁优雅的实现方式,既能快速执行、易于维护,又不会造成代码冗余和命名空间污染?


方案1:内部匿名函数列表+循环调用

把提取逻辑用匿名函数(lambda)定义在get_date函数内部,既不会污染外部命名空间,又能清晰罗列所有提取规则,循环执行直到找到有效日期:

def get_date(html):
    # 内部定义所有提取规则,不污染外部命名空间
    extractors = [
        lambda h: h.find("time", {"datetime": True})["datetime"] if h.find("time", {"datetime": True}) else None,
        lambda h: h.find("meta", {"property": "article:modified_time", "content": True})["content"] if h.find("meta", {"property": "article:modified_time", "content": True}) else None,
        lambda h: h.find("div", {"class": "dateline"}).get_text() if h.find("div", {"class": "dateline"}) else None
        # 后续规则直接追加到列表即可
    ]
    
    for extractor in extractors:
        date = extractor(html)
        if date:
            return date
    return None

方案2:生成器表达式+next() 简化循环

利用Python的next()函数可以接收默认值的特性,用生成器表达式遍历提取规则,找到第一个非None的结果就返回,代码更紧凑:

def get_date(html):
    extractors = [
        lambda h: h.find("time", {"datetime": True})["datetime"] if h.find("time", {"datetime": True}) else None,
        lambda h: h.find("meta", {"property": "article:modified_time", "content": True})["content"] if h.find("meta", {"property": "article:modified_time", "content": True}) else None,
        lambda h: h.find("div", {"class": "dateline"}).get_text() if h.find("div", {"class": "dateline"}) else None
    ]
    
    # 生成器逐个执行提取器,返回第一个非None的结果,无结果则返回None
    return next((extractor(html) for extractor in extractors if extractor(html)), None)

方案3:提取规则结构化(最易维护)

把每个提取规则拆分为「查找参数」和「取值逻辑」的结构化数据,比如元组,这样后期修改、新增规则时不需要写lambda,直接修改规则列表即可,可读性和维护性更强:

def get_date(html):
    # 每个规则是一个元组:(查找标签, 查找属性, 取值函数)
    extraction_rules = [
        ("time", {"datetime": True}, lambda tag: tag["datetime"]),
        ("meta", {"property": "article:modified_time", "content": True}, lambda tag: tag["content"]),
        ("div", {"class": "dateline"}, lambda tag: tag.get_text())
        # 新增规则只需按格式追加
    ]
    
    for tag_name, attrs, getter in extraction_rules:
        tag = html.find(tag_name, attrs)
        if tag:
            try:
                return getter(tag)
            except (KeyError, AttributeError):
                # 防止标签存在但取值失败的情况,继续尝试下一个规则
                continue
    return None

这种方式的优势在于:规则和逻辑完全分离,新增或修改规则时只需要调整extraction_rules列表,不需要改动循环逻辑;同时通过异常捕获处理潜在的取值错误,鲁棒性更强。


内容的提问来源于stack exchange,提问作者Aurélien Pierre

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 11:42:25