You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提取无HTML标签文本并设置split多分隔符?多分割点提取div内容

Answers to Your Web Scraping Questions

Question 1: Extracting Text Without HTML Tags & Using Multiple Separators in Split

Extracting Text Without HTML Tags

Since you’re using BeautifulSoup (evident from your second question), here are two reliable ways to pull clean text free of HTML tags:

  • Use .get_text(): This method aggregates all text from the element and its children, with options to control whitespace and separators. Example:
    clean_text = div.get_text(separator=' ', strip=True)
    
    strip=True removes leading/trailing whitespace, while separator=' ' ensures consistent spacing between text blocks instead of messy newlines or tabs.
  • Use .stripped_strings: This returns a generator that yields text chunks with all extra whitespace stripped out. You can join them into a single clean string like this:
    clean_text = " ".join(div.stripped_strings)
    
    This is exactly what you’re already using in your code—great pick for consistent whitespace handling!

Splitting with Multiple Separators

Python’s built-in str.split() only supports one fixed separator (or a set of characters as a single separator). To split on multiple distinct strings, use re.split() from the re module, which lets you define a regex pattern matching any of your target separators.

For example, splitting on "cat" or "dog":

import re
text = "My cat and dog play together"
parts = re.split(r'cat|dog', text)
# Result: ['My ', ' and ', ' play together']

Note: If your separators contain regex special characters (like ., *), escape them with a backslash (e.g., r'\.' for a literal dot).


Question 2: Handling Multiple Split Points for Prerequisites/Corerequisites

Your current code only splits on "Prerequisite: ", but we can expand this to match all four possible split points using a regex pattern with re.split(). Here’s the adjusted code:

import re
from bs4 import BeautifulSoup

# Assume you've already parsed your HTML into the `soup` object
div = soup.select("div.ajaxcourseindentfix")[0]
clean_text = " ".join(div.stripped_strings)

# Regex pattern to match all four split variations
# Matches: Prerequisite:, Prerequisites:, Corerequisite:, Corerequisites:
split_pattern = r'(Prerequisite(s)?|Corerequisite(s)?):\s*'

# Split the text and grab everything after the last match
result = re.split(split_pattern, clean_text)[-1].strip()

How the Pattern Works:

  • (Prerequisite(s)?|Corerequisite(s)?):: This matches either Prerequisite:/Prerequisites: or Corerequisite:/Corerequisites—the (s)? makes the trailing "s" optional.
  • \s*: Matches any number of whitespace characters after the colon, so we don’t end up with leading spaces in our final result.

Edge Case Handling:

If none of the split points are found, re.split() will return a list with just the original clean_text. To handle this (e.g., return an empty string or a message), add a quick check:

split_parts = re.split(split_pattern, clean_text)
if len(split_parts) > 1:
    result = split_parts[-1].strip()
else:
    result = ""  # Or "No prerequisites/corerequisites found"

内容的提问来源于stack exchange,提问作者miserable

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 11:35:20