使用Python和BeautifulSoup爬取日期遇分割问题求助
解决BeautifulSoup爬取日期时的内容提取问题
你的问题出在两个核心点:
- 目标
<p>标签没有datetime属性,所以get('datetime')自然无法获取到值; - 你误用了
.作为分割符,但原文本里日期和阅读时长的分隔符是•(中间圆点),同时split的参数用法也不符合语法要求。
下面是两种可靠的提取方案:
方法1:按分隔符分割文本
先获取标签的纯文本内容,再用•分割取前半部分:
from bs4 import BeautifulSoup # 示例HTML,实际替换为你爬取的页面内容 html = '<p class="text-xs">Oct 24, 2017 • 4 min read</p>' soup = BeautifulSoup(html, 'html.parser') # 获取标签文本并清理首尾空白 text = soup.select_one('p.text-xs').get_text(strip=True) # 按中间圆点分割,取第一个部分再清理空白 published_date = text.split('•')[0].strip() print(published_date) # 输出: Oct 24, 2017
方法2:正则表达式匹配(更通用)
如果日期格式固定(英文月份缩写+日期+四位年份),用正则匹配能避免分隔符变化带来的问题:
import re from bs4 import BeautifulSoup html = '<p class="text-xs">Oct 24, 2017 • 4 min read</p>' soup = BeautifulSoup(html, 'html.parser') text = soup.select_one('p.text-xs').get_text(strip=True) # 匹配 "英文月份缩写 数字, 四位年份" 的格式 date_match = re.search(r'([A-Za-z]{3} \d{1,2}, \d{4})', text) if date_match: published_date = date_match.group(1) print(published_date) # 输出: Oct 24, 2017
内容的提问来源于stack exchange,提问作者Info Rewind
相关产品推荐
相关产品推荐

