You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用正则表达式提取<td></td>内详情?Cvedetails表格数据提取求助

用正则表达式提取Cvedetails表格数据及标签内容

Hey there, let's break down how to tackle this regex extraction for Cvedetails. First off, let's start with the basics of pulling content out of <td> tags, then move up to grabbing entire table rows.

一、提取单个标签内的内容

For getting the content inside a <td> tag, here's a solid base regex pattern:

<td[^>]*>(.*?)</td>

Let me break this down so you understand what's happening:

  • <td[^>]*>: Matches the opening <td tag, including any attributes like class or align (the [^>]* grabs everything until the closing > of the tag).
  • (.*?): Uses non-greedy matching to capture everything inside the tag—this is crucial because it stops at the first </td> instead of matching all the way to the last one in the table.
  • </td>: Matches the closing tag to wrap up the pattern.

If you need to target a specific <td> with a class (like those labeled cve-number on Cvedetails), tweak the regex to lock onto that attribute:

<td class="cve-number"[^>]*>(.*?)</td>

二、提取整行表格数据

Cvedetails tables use <tr> tags to wrap each row. To pull all the data from a row, you'll first match the entire row, then extract each <td> inside it. Here's a practical example using Python:

import re

# Sample HTML snippet from Cvedetails
html_snippet = """
<tr class="srrowns">
    <td>CVE-2024-5678</td>
    <td><span class="cvssboxcritical">Critical</span></td>
    <td>2024-06-10</td>
</tr>
"""

# First, grab all table rows
rows = re.findall(r'<tr[^>]*>(.*?)</tr>', html_snippet, re.DOTALL)

for row in rows:
    # Extract each cell's content from the row
    cells = re.findall(r'<td[^>]*>(.*?)</td>', row, re.DOTALL)
    # Clean up the content (strip extra tags and whitespace)
    cleaned_cells = [re.sub(r'<.*?>', '', cell).strip() for cell in cells]
    print("Extracted row data:", cleaned_cells)

The re.DOTALL flag is key here—it makes the . character match newlines, which is necessary because HTML rows are usually split across multiple lines.

重要提醒

While regex works for simple cases, HTML is not a regular language, so regex can break if the page structure changes (like nested tags, hidden elements, or weird formatting). If you're planning on scaling this or dealing with more complex tables, I'd strongly suggest using an HTML parser like BeautifulSoup (Python) or Cheerio (JavaScript)—they handle edge cases way better than regex ever will. But if regex is your only tool right now, test your patterns against different parts of the Cvedetails page to make sure they hold up.

内容的提问来源于stack exchange,提问作者Bryead

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 07:36:06