如何获取网页文本?详解JavaScript方法与BeautifulSoup的get_text()函数
get_text() Great question! Let's break this down into two clear parts—grabbing webpage text with JavaScript, and using Python's BeautifulSoup to extract text from parsed HTML.
JavaScript: Fetching Webpage Text
There are a couple of go-to methods to pull text from a webpage or specific elements with JS:
Get all text in the body (including hidden elements): Use
textContent—it grabs every piece of text inside the body, even from elements that are hidden via CSS.const fullPageText = document.body.textContent;Get only visible text: If you want text that users can actually see on the page, use
innerTextinstead. It respects CSS visibility rules.const visiblePageText = document.body.innerText;Extract text from a specific element: For a target element (like the
<span>you mentioned), target it by ID, class, or selector, then usetextContent:// Extract text from the span with ID "sm_flash_225" const spanText = document.getElementById('sm_flash_225').textContent; // This will return "bail" for your example span
Python: Using BeautifulSoup's get_text()
BeautifulSoup's get_text() is perfect for stripping all HTML tags and pulling out pure text from parsed HTML content. Here's how it works:
Basic Usage
First, parse your HTML into a BeautifulSoup object (you'll usually fetch the page with requests first, but you can also pass raw HTML strings):
from bs4 import BeautifulSoup import requests # Example: Fetch and parse a webpage url = "your-webpage-url-here" response = requests.get(url) soup = BeautifulSoup(response.text, 'html.parser') # Extract all text from the parsed page full_text = soup.get_text()
Extract Text from a Specific Element
For the <span> example you provided, it's just as simple:
html = '<span id="sm_flash_225" onclick="sm_flash_process(\'bail\', this,1)" onmouseover="sm_flash_add(\'bail\', this, 1);" onmouseout="sm_flash_remove(\'bail\', this, 1);">bail</span>' soup = BeautifulSoup(html, 'html.parser') span_text = soup.get_text() print(span_text) # Output: bail
Customize the Output
get_text() has useful parameters to clean up the text:
- Use a separator (like
\n) to split text from different elements:formatted_text = soup.get_text('\n') - Add
strip=Trueto remove extra whitespace, newlines, and tabs from the start/end of each text segment:clean_text = soup.get_text('\n', strip=True)
When you run this on a full parsed HTML page, it will strip all tags and return a single block of pure text—exactly like the example result you referenced.
内容的提问来源于stack exchange,提问作者Foobar

