You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python使用Selenium如何定位类名可变的论坛文章发布时间元素

Selenium同时匹配普通文章和megapost发布时间解决方案

问题背景

我正在从论坛抓取普通文章和megapost的指标数据,包括点赞量、浏览量、评论数、发布日期。
我使用Selenium工具提取内容的发布时间,现有实现代码如下:

dates = []

page_items = len(drv.find_elements_by_class_name("tm-articles-list"))
for i in range(page_items):
    date_of_post = drv.find_elements_by_class_name("tm-article-snippet__datetime-published")
    for d in date_of_post:
        date_text = d.find_element_by_tag_name("time").text
        dates.append(date_text)

遇到的问题

普通文章和megapost的发布时间类名存在差异:

  • 普通文章对应类名:tm-article-snippet__datetime-published
  • megapost对应类名:tm-megapost-snippet__datetime-published
    尝试用Python逻辑运算符or拼接类名查询失效,需要找到不区分类别即可解析所有发布时间的方法。
    所有文章(含megapost)都位于tm-articles-list类的节点下。

解决代码

直接使用CSS选择器的多规则匹配特性即可实现需求:

dates = []
# 逗号分隔多个选择规则,满足任意规则就会被匹配
date_elements = drv.find_elements_by_css_selector(
    ".tm-article-snippet__datetime-published, .tm-megapost-snippet__datetime-published"
)
for ele in date_elements:
    date_text = ele.find_element_by_tag_name("time").text
    dates.append(date_text)

如果需要限定仅在文章列表范围内匹配,避免匹配到页面其他区域的时间节点,可以调整选择器为:

date_elements = drv.find_elements_by_css_selector(
    ".tm-articles-list .tm-article-snippet__datetime-published, .tm-articles-list .tm-megapost-snippet__datetime-published"
)

原错误写法原因说明

Python的or是逻辑运算符,对两个非空字符串做逻辑判断时,会直接返回第一个为真的值,也就是你传入find_elements_by_class_name的参数实际只有tm-article-snippet__datetime-published,自然匹配不到megapost的时间节点。


内容的提问来源于stack exchange,提问作者rg4s

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 11:21:03