You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用BeautifulSoup4查找无class属性的<p>标签?

问题:从MITRE ATT&CK页面提取无class属性的

标签内容

我正在做Python项目提升技能,需要从https://attack.mitre.org/techniques/T1588/001/这类页面中提取无class属性的

标签内的数据,用BeautifulSoup解析。

现有代码及问题

我当前的代码片段如下:

import requests
from bs4 import BeautifulSoup
import csv
import re

mitigation = []

for id in id_list:
    page = requests.get(f"https://attack.mitre.org/techniques/{id2}")
    soup = BeautifulSoup(page.content, 'html.parser')
    paragraph = soup.find('p', class_ = '')
    status_code = page.status_code
    mitigation.append(paragraph)

用soup.find('p', class_ = '')没法正确获取目标内容,后来试过用固定索引定位:

paragraph = soup.select('p')[4].text

但这种方式不稳定,页面结构一旦变化就会失效,希望找到可靠的提取方法。


解决方案

针对MITRE ATT&CK页面的结构,推荐几种精准提取的方法:

1. CSS选择器筛选无class的p标签

用:not()伪类排除带class属性的p标签,再结合父容器缩小范围(比如技术描述区域):

import requests
from bs4 import BeautifulSoup
import csv
import re

mitigation = []

for id in id_list:
    # 修正变量名笔误:id2 → id
    page = requests.get(f"https://attack.mitre.org/techniques/{id}")
    if page.status_code != 200:
        print(f"请求失败:{id}")
        continue
    soup = BeautifulSoup(page.content, 'html.parser')
    # 定位技术描述的父容器
    desc_container = soup.find('div', class_='technique-description')
    if desc_container:
        # 筛选容器内无class的p标签
        target_paragraphs = desc_container.select('p:not([class])')
        # 提取第一个非空的目标内容
        for p in target_paragraphs:
            text = p.get_text(strip=True)
            if text:
                mitigation.append(text)
                break

2. 遍历判断p标签的class属性

遍历所有p标签,检查是否无class属性:

for p in soup.find_all('p'):
    # 判定条件:无class属性 或 class为空列表/空字符串
    if not p.get('class'):
        text = p.get_text(strip=True)
        if text:
            mitigation.append(text)
            break  # 仅提取第一个符合条件的内容则break

3. 结合页面标题定位

MITRE页面的目标描述通常在"Description"标题之后,可先找到对应标题再取后续的p标签:

desc_heading = soup.find('h2', string=re.compile('Description', re.IGNORECASE))
if desc_heading:
    target_p = desc_heading.find_next_sibling('p')
    if target_p and not target_p.get('class'):
        mitigation.append(target_p.get_text(strip=True))

额外提示

  • 代码中id2是笔误,需改为循环变量id
  • 务必加入请求状态码判断,避免解析失败页面
  • 避免用固定索引(如[4])定位,页面结构更新后会直接失效

内容的提问来源于stack exchange,提问作者TrashThrash

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 21:22:47