如何使用Python或Selenium提取HTML中指定div内的BURGUNDY文本
Got it, let's figure out how to extract that "BURGUNDY" text! The tricky part here is that your target text is a standalone text node sitting right inside the .row div, right after the empty .col-md-12 div. A lot of basic text extraction methods might skip over it because they only look at text directly attached to elements, not sibling text nodes. Here are two reliable solutions using Python tools:
Solution 1: Using BeautifulSoup (Static HTML Parsing)
If you're working with static HTML (not a live browser page), BeautifulSoup is perfect for this. First, install the package if you haven't:
pip install beautifulsoup4
Then use this code to parse and extract the text:
from bs4 import BeautifulSoup # Your raw HTML html_content = ''' <div class = "row"><div class = "col-md-12"></div>BURGUNDY</div> <div class = "row"><div class = "col-md-12"></div>randomTxt</div> ''' # Parse the HTML soup = BeautifulSoup(html_content, 'html.parser') # Target the first .row element first_row = soup.find('div', class_='row') # Get all stripped text strings from the row, filter out empty ones clean_texts = [text.strip() for text in first_row.stripped_strings if text.strip()] # The "BURGUNDY" will be the first non-empty entry burgundy_text = clean_texts[0] if clean_texts else None print(burgundy_text) # Output: BURGUNDY
The stripped_strings method iterates over every text node inside the .row element, automatically stripping whitespace. We just filter out any empty strings left from the empty .col-md-12 div, and we're left with our target text.
Solution 2: Using Selenium (Live Browser Sessions)
If you're dealing with a dynamic page (where the HTML loads after page load), Selenium can access the text nodes via JavaScript execution. Here's how:
First, install Selenium and set up your browser driver (e.g., ChromeDriver):
pip install selenium
Then use this code:
from selenium import webdriver from selenium.webdriver.common.by import By # Initialize the browser driver (use the one matching your browser) driver = webdriver.Chrome() # For testing, we'll inject your HTML into the browser test_html = ''' <div class = "row"><div class = "col-md-12"></div>BURGUNDY</div> <div class = "row"><div class = "col-md-12"></div>randomTxt</div> ''' driver.execute_script(f"document.body.innerHTML = '{test_html}';") # Find the first .row element target_row = driver.find_element(By.CLASS_NAME, 'row') # Use JavaScript to extract the text node burgundy_text = driver.execute_script(""" const rowElement = arguments[0]; // Get all child nodes, filter to only text nodes const textNodes = Array.from(rowElement.childNodes).filter(node => node.nodeType === 3); // Trim text and filter out empty strings, then grab the first valid one return textNodes.map(node => node.textContent.trim()).filter(txt => txt)[0]; """, target_row) print(burgundy_text) # Output: BURGUNDY # Don't forget to close the driver driver.quit()
This works because we're directly accessing the DOM's text nodes using JavaScript. We filter nodes by nodeType === 3 (which identifies text nodes), then clean up the text to get our target value.
If you need to extract text from all .row elements instead of just the first, you can loop through driver.find_elements(By.CLASS_NAME, 'row') and repeat the extraction for each one.
内容的提问来源于stack exchange,提问作者Fity

