You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python和BeautifulSoup爬取日期遇分割问题求助

解决BeautifulSoup爬取日期时的内容提取问题

你的问题出在两个核心点:

  1. 目标<p>标签没有datetime属性,所以get('datetime')自然无法获取到值;
  2. 你误用了.作为分割符,但原文本里日期和阅读时长的分隔符是•(中间圆点),同时split的参数用法也不符合语法要求。

下面是两种可靠的提取方案:

方法1:按分隔符分割文本

先获取标签的纯文本内容,再用•分割取前半部分:

from bs4 import BeautifulSoup

# 示例HTML,实际替换为你爬取的页面内容
html = '<p class="text-xs">Oct 24, 2017 • 4 min read</p>'
soup = BeautifulSoup(html, 'html.parser')

# 获取标签文本并清理首尾空白
text = soup.select_one('p.text-xs').get_text(strip=True)
# 按中间圆点分割,取第一个部分再清理空白
published_date = text.split('•')[0].strip()
print(published_date)  # 输出: Oct 24, 2017

方法2:正则表达式匹配(更通用)

如果日期格式固定(英文月份缩写+日期+四位年份),用正则匹配能避免分隔符变化带来的问题:

import re
from bs4 import BeautifulSoup

html = '<p class="text-xs">Oct 24, 2017 • 4 min read</p>'
soup = BeautifulSoup(html, 'html.parser')

text = soup.select_one('p.text-xs').get_text(strip=True)
# 匹配 "英文月份缩写 数字, 四位年份" 的格式
date_match = re.search(r'([A-Za-z]{3} \d{1,2}, \d{4})', text)
if date_match:
    published_date = date_match.group(1)
    print(published_date)  # 输出: Oct 24, 2017

内容的提问来源于stack exchange,提问作者Info Rewind

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 14:10:28