You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何获取网页文本?详解JavaScript方法与BeautifulSoup的get_text()函数

How to Extract Text from Webpages: JavaScript & BeautifulSoup's get_text()

Great question! Let's break this down into two clear parts—grabbing webpage text with JavaScript, and using Python's BeautifulSoup to extract text from parsed HTML.

JavaScript: Fetching Webpage Text

There are a couple of go-to methods to pull text from a webpage or specific elements with JS:

  • Get all text in the body (including hidden elements): Use textContent—it grabs every piece of text inside the body, even from elements that are hidden via CSS.

    const fullPageText = document.body.textContent;
    
  • Get only visible text: If you want text that users can actually see on the page, use innerText instead. It respects CSS visibility rules.

    const visiblePageText = document.body.innerText;
    
  • Extract text from a specific element: For a target element (like the <span> you mentioned), target it by ID, class, or selector, then use textContent:

    // Extract text from the span with ID "sm_flash_225"
    const spanText = document.getElementById('sm_flash_225').textContent;
    // This will return "bail" for your example span
    

Python: Using BeautifulSoup's get_text()

BeautifulSoup's get_text() is perfect for stripping all HTML tags and pulling out pure text from parsed HTML content. Here's how it works:

Basic Usage

First, parse your HTML into a BeautifulSoup object (you'll usually fetch the page with requests first, but you can also pass raw HTML strings):

from bs4 import BeautifulSoup
import requests

# Example: Fetch and parse a webpage
url = "your-webpage-url-here"
response = requests.get(url)
soup = BeautifulSoup(response.text, 'html.parser')

# Extract all text from the parsed page
full_text = soup.get_text()

Extract Text from a Specific Element

For the <span> example you provided, it's just as simple:

html = '<span id="sm_flash_225" onclick="sm_flash_process(\'bail\', this,1)" onmouseover="sm_flash_add(\'bail\', this, 1);" onmouseout="sm_flash_remove(\'bail\', this, 1);">bail</span>'
soup = BeautifulSoup(html, 'html.parser')
span_text = soup.get_text()
print(span_text)  # Output: bail

Customize the Output

get_text() has useful parameters to clean up the text:

  • Use a separator (like \n) to split text from different elements:
    formatted_text = soup.get_text('\n')
    
  • Add strip=True to remove extra whitespace, newlines, and tabs from the start/end of each text segment:
    clean_text = soup.get_text('\n', strip=True)
    

When you run this on a full parsed HTML page, it will strip all tags and return a single block of pure text—exactly like the example result you referenced.

内容的提问来源于stack exchange,提问作者Foobar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:57:51