You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python遍历字符串提取所有datetime对应日期值的问题

问题:从HTML字符串中提取所有datetime属性值

现有一段长HTML字符串示例:

<relative-time class="no-wrap" datetime="2023-03-07T02:38:29Z" title="Mar 6, 2023, 7:38 PM MST">Mar 6, 2023</relative-time>, <relative-time data-view-component="true" datetime="2023-03-06T10:25:38-07:00

需要提取每个datetime属性对应的日期值。当前实现思路是先获取所有datetime的索引:

datetime_indexes = list(get_all_updates(string, "datetime"))
print(datetime_indexes)
# output 36, 168 etc

再通过循环匹配索引提取日期:

count = 0
all_datetimes = []
for i in string:
    if string.index(i) is datetime_indexes[count]:
        all_datetimes.append(string[string.index(i) + 10:(string.index(i) + 10 + 21)])
        count = count + 1

但仅能输出第一个datetime值:

# output
#2023-03-07T02:38:29Z

期望结果是获取所有datetime值:

# desired output
2023-03-07T02:38:29
2023-03-06T10:25:38

解决方案

1. 修复索引遍历逻辑

当前代码的核心问题:

  • string.index(i)会返回字符i第一次出现的索引,而非当前遍历位置,导致匹配逻辑失效
  • 逐个遍历字符效率极低,完全没必要

直接遍历已获取的datetime_indexes即可:

all_datetimes = []
# 假设get_all_updates能正确返回所有"datetime"子串的起始索引
for idx in datetime_indexes:
    # 定位到datetime属性值的起始位置(跳过"datetime=\"")
    start = idx + len('datetime="')
    # 找到属性值的结束引号位置
    end = string.find('"', start)
    if end != -1:
        dt_value = string[start:end]
        # 去掉末尾的Z或时区偏移,保留到秒级格式
        dt_value = dt_value.split('Z')[0].split('-')[0].split('+')[0]
        all_datetimes.append(dt_value)

print(all_datetimes)
# 输出: ['2023-03-07T02:38:29', '2023-03-06T10:25:38']

2. 一步到位:用正则表达式

直接匹配所有datetime="..."中的内容,无需手动处理索引:

import re

html_str = '<relative-time class="no-wrap" datetime="2023-03-07T02:38:29Z" title="Mar 6, 2023, 7:38 PM MST">Mar 6, 2023</relative-time>, <relative-time data-view-component="true" datetime="2023-03-06T10:25:38-07:00'

# 匹配datetime属性值,捕获引号内的内容
pattern = r'datetime="([^"]+)"'
matches = re.findall(pattern, html_str)

# 清理格式,去掉时区部分
cleaned_dts = [dt.split('Z')[0].split('-')[0].split('+')[0] for dt in matches]

print(cleaned_dts)
# 输出: ['2023-03-07T02:38:29', '2023-03-06T10:25:38']

3. 最可靠方案:用HTML解析库(推荐)

处理HTML优先用专门的解析库,避免正则的局限性(比如属性值用单引号的场景):

from bs4 import BeautifulSoup

html_str = '<relative-time class="no-wrap" datetime="2023-03-07T02:38:29Z" title="Mar 6, 2023, 7:38 PM MST">Mar 6, 2023</relative-time>, <relative-time data-view-component="true" datetime="2023-03-06T10:25:38-07:00'

soup = BeautifulSoup(html_str, 'html.parser')
# 筛选所有带有datetime属性的标签
target_tags = soup.find_all(attrs={"datetime": True})

all_datetimes = []
for tag in target_tags:
    dt_value = tag['datetime']
    # 清理格式
    cleaned_dt = dt_value.split('Z')[0].split('-')[0].split('+')[0]
    all_datetimes.append(cleaned_dt)

print(all_datetimes)
# 输出: ['2023-03-07T02:38:29', '2023-03-06T10:25:38']

内容的提问来源于stack exchange,提问作者TripleCute

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 09:53:04