You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用正则在多个HTML页面中匹配识别多种格式的日期

多日期格式匹配修改方案

你只需将多个正则规则用正则的或运算符|拼接为一个复合规则,即可一次性匹配所有符合要求的日期格式。同时注意修正原正则中的特殊字符转义问题,删除冗余代码。

修改后的完整代码

from selenium import webdriver
from bs4 import BeautifulSoup
import re

headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/92.0.4515.131 Safari/537.3"
}

url_list = ['https://www.grsymphony.org/american-in-paris','https://www.toledosymphony.com/events/2021/06/04/promedica-masterworks/beethoven-s-pastoral/',
           'https://www.toledosymphony.com/events/2021/09/24/promedica-masterworks/musical-portraits/',
           'https://www.staatsoper.de/stueckinfo/les-troyens-1/2022-05-09-17-00.html?tx_sfstaatsoper_pi1%5BfromSpielplan%5D=1&tx_sfstaatsoper_pi1%5BpageId%5D=528&cHash=3ce5142af1140c90522372caa4330efb',
           'https://www.hso.org/concerts/an-innocent-man-the-music-of-billy-joel/',
           'https://www.hso.org/concerts/beethoven-triple-concerto/','https://www.seattlesymphony.org/en/concerttickets/calendar/2021-2022/21bar3',]

driver = webdriver.Chrome('/home/ubuntu/selenium_drivers/chromedriver')

# 合并所有日期匹配规则,注意第一个规则的点号做转义处理
date_pattern = r'[ADFJMNOS]\w* [\d]{1,2}\. [\d]{4}|[\d]{1,2} [ADFJMNOS]\w*, [\d]{4}|[\d]{4}, [ADFJMNOS]\w* [\d]{1,2}|[ADFJMNOS]\w* [\d]{1,2} - [\d]{1,2}, [\d]{4}'

for URL in url_list:
    driver.get(URL)
    driver.implicitly_wait(2)
    data = driver.page_source
    # BeautifulSoup的text属性已经自动剥离HTML标签,不需要额外做正则替换
    cleantext = BeautifulSoup(data, "lxml").text
    all_dates = re.findall(date_pattern, cleantext) 
    print(URL)
    for s in all_dates:
        print(s)

核心修改说明

  • 复合正则规则用|分隔不同的匹配模式,re.findall会自动返回所有符合任意一种模式的结果
  • 对第一个规则里的.做了转义处理(改为\.),避免正则把点识别为通配符匹配到无关内容
  • 删除了冗余的HTML标签清理代码,减少不必要的性能开销

内容的提问来源于stack exchange,提问作者Aniiya0978

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 10:51:03