You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup爬取含唯一后缀类名的牙科论坛主题类名列表

解决方案

修改后的代码

from bs4 import BeautifulSoup
import requests

url = "https://www.dentistry-forums.com/forums/periodontics.11/"
result = requests.get(url).text
doc = BeautifulSoup(result, "html.parser")

# 直接匹配所有带目标动态后缀类名的主题元素
thread_items = doc.select('div[class*="js-threadListItem-"]')

# 收集所有元素的完整类名字符串
class_name_list = []
for item in thread_items:
    # 将class列表转为空格分隔的完整字符串,匹配示例格式
    full_class_str = ' '.join(item['class'])
    class_name_list.append(full_class_str)

# 输出结果
print(class_name_list)

关键说明

  • 简化定位逻辑:替换逐层find的冗余写法,用BeautifulSoup的select方法结合CSS属性选择器[class*="js-threadListItem-"],直接命中所有带动态后缀类名的元素,代码更简洁高效。
  • 精准匹配备选方案:如果担心误匹配其他元素,可结合固定类名缩小范围,确保只抓取主题项:
    thread_items = doc.select('div.structItem.structItem--thread[class*="js-threadListItem-"]')
    
  • 类名格式转换:BeautifulSoup返回的元素class属性是列表类型,通过' '.join()将其转为空格分隔的完整字符串,和你给出的示例格式完全一致。

内容的提问来源于stack exchange,提问作者jack gell

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 02:05:17