如何用Python+Beautiful Soup/Selenium提取HTML指定子元素文本
提取HTML中的三段文本内容
下面提供两种常用工具的解决方案,分别适用于静态HTML解析的Beautiful Soup和动态页面处理的Selenium:
使用Beautiful Soup(静态HTML)
方法1:利用stripped_strings提取有序文本节点
stripped_strings会返回当前元素下所有去除空白的文本内容,按文档顺序排列,直接取第一个即可得到目标第一段文本:
from bs4 import BeautifulSoup html = ''' <div class="price-options "> Was $8.50, Save $4.25 <br> <span class="price"> $4.25 each</span> <br> <span class="comparative-text">$1.12 per 100g</span> </div> ''' soup = BeautifulSoup(html, 'html.parser') price_div = soup.find('div', class_='price-options') # 提取第一段文本 first_text = next(price_div.stripped_strings) # 提取后两段已知文本 second_text = price_div.find('span', class_='price').get_text(strip=True) third_text = price_div.find('span', class_='comparative-text').get_text(strip=True) print(first_text) # 输出: Was $8.50, Save $4.25 print(second_text) # 输出: $4.25 each print(third_text) # 输出: $1.12 per 100g
方法2:遍历子节点筛选文本
直接遍历div的子节点,找到第一个非空白的文本节点:
first_text = "" for child in price_div.children: # 判断是否为文本节点且内容非空 if isinstance(child, str) and child.strip(): first_text = child.strip() break
使用Selenium(动态页面)
如果页面内容是动态渲染的,用Selenium获取元素文本后,按换行分割并过滤空行即可:
from selenium import webdriver from selenium.webdriver.common.by import By driver = webdriver.Chrome() driver.get("目标页面URL") # 提取第一段文本 price_div = driver.find_element(By.CLASS_NAME, "price-options") all_lines = price_div.text.split('\n') first_text = next(line.strip() for line in all_lines if line.strip()) # 提取后两段文本 second_text = driver.find_element(By.CLASS_NAME, "price").text.strip() third_text = driver.find_element(By.CLASS_NAME, "comparative-text").text.strip() print(first_text) print(second_text) print(third_text) driver.quit()
方案对比
- 直接提取文本节点的方式比拆分总文本或移除已提取部分更简洁,避免了因HTML结构变化导致的拆分错误。
- 静态HTML优先用Beautiful Soup,动态页面则选择Selenium。
内容的提问来源于stack exchange,提问作者Hammertime
相关产品推荐
相关产品推荐

