You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从GET返回的HTML字符串中提取指定标签及类的内容?

Extracting Specific Content from an HTML String

Hey there! Great question—pulling out content tied to a specific tag and class from raw HTML is a routine task, and there are solid approaches depending on what environment you're working in. Here are the most reliable methods:

1. Browser/Client-Side JavaScript

If you're working in a browser, you can leverage the native DOM API to parse the HTML string and target your desired element:

// Your raw HTML string from the GET request
const htmlString = '<div class="header">...</div><article class="content">This is the content I want!</article><footer>...</footer>';

// Create a temporary DOM element to hold the HTML
const tempDiv = document.createElement('div');
tempDiv.innerHTML = htmlString;

// Use querySelector to target the element by tag + class (adjust selector as needed)
const contentElement = tempDiv.querySelector('article.content'); // Or just '.content' if class is unique

// Extract the content—use innerHTML to keep tags, or textContent for plain text
const extractedContent = contentElement ? contentElement.innerHTML : 'No content found';
console.log(extractedContent);

2. Node.js Environment

For server-side JavaScript, use Cheerio—a lightweight library that mimics jQuery's syntax for parsing HTML:
First install Cheerio:

npm install cheerio

Then use it to extract your content:

const cheerio = require('cheerio');

const htmlString = '<div class="header">...</div><section class="content"><p>Here’s the targeted content!</p></section>';
const $ = cheerio.load(htmlString);

// Target the element with your selector
const extractedContent = $('section.content').html(); // Or .text() for plain text
console.log(extractedContent);

3. Python Environment

In Python, Beautiful Soup is the go-to tool for HTML parsing:
First install Beautiful Soup and a parser (like lxml):

pip install beautifulsoup4 lxml

Then parse and extract:

from bs4 import BeautifulSoup

html_string = '<div class="sidebar">...</div><div class="content"><h2>My Content</h2><p>Hello world!</p></div>'
soup = BeautifulSoup(html_string, 'lxml')

# Find the element by tag and class—use find() for the first match, find_all() for multiple
content_element = soup.find('div', class_='content')

if content_element:
    extracted_content = content_element.prettify() # Or .text for plain text
    print(extracted_content)
else:
    print("No content element found")

Critical Note

Never use regular expressions to parse HTML—HTML is a nested, non-regular language, and regex will fail on edge cases like nested tags, unexpected whitespace, or reordered attributes. Always use a dedicated HTML parser like the ones above.

内容的提问来源于stack exchange,提问作者Dextranovich

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:04:38