新手求助:如何从div标签内的span标签中提取数据?
Hey there! No worries at all—everyone starts with basic questions when diving into web scraping. Let's walk through exactly how to pull the data you need from that HTML snippet.
提取目标数据的分步指南
We'll use Python's BeautifulSoup library for this—it's super beginner-friendly and perfect for parsing static HTML like your example.
Step 1: Set up your environment
First, make sure you have the necessary libraries installed:
pip install beautifulsoup4 requests
(We'll use requests only if you're pulling HTML directly from a live website; if you already have the HTML snippet, you can skip the requests part.)
Step 2: Parse HTML and extract data
Here's a complete code example tailored to your HTML structure (I filled in the truncated old price part for clarity):
from bs4 import BeautifulSoup # Your HTML snippet (with the truncated old price completed) html_content = ''' <div class="price-container clearfix"> <span class="sale-flag-percent">-40%</span> <span class="price-box ri"> <span class="price "><span data-currency-iso="PKR">Rs.</span> <span dir="ltr" data-price="5999"> 5,999</span> </span> <span class="price -old "><span data-currency-iso="PKR">Rs.</span> <span dir="ltr" data-price="9999"> 9,999</span></span> </span> </div> ''' # Parse the HTML soup = BeautifulSoup(html_content, 'html.parser') # Locate the main price container div price_container = soup.find('div', class_='price-container clearfix') # Extract discount percentage discount = price_container.find('span', class_='sale-flag-percent').text.strip() print(f"Discount: {discount}") # Extract current price details current_price_section = price_container.find('span', class_='price ') currency_symbol = current_price_section.find('span', attrs={'data-currency-iso': 'PKR'}).text.strip() current_price_text = current_price_section.find('span', attrs={'data-price': True}).text.strip() current_price_raw = current_price_section.find('span', attrs={'data-price': True})['data-price'] # Get pure numeric value print(f"Current Price: {currency_symbol}{current_price_text} (Raw value: {current_price_raw})") # Extract old price details old_price_section = price_container.find('span', class_='price -old ') old_currency_symbol = old_price_section.find('span', attrs={'data-currency-iso': 'PKR'}).text.strip() old_price_text = old_price_section.find('span', attrs={'data-price': True}).text.strip() old_price_raw = old_price_section.find('span', attrs={'data-price': True})['data-price'] print(f"Original Price: {old_currency_symbol}{old_price_text} (Raw value: {old_price_raw})")
Key explanations
soup.find(): Locates a single HTML element using attributes likeclass_or custom attributes (likedata-currency-iso).text.strip(): Pulls the visible text inside a tag and cleans up extra spaces/newlines['data-price']: Directly accesses the value of thedata-priceattribute, which gives you a clean numeric price (great for calculations later)
Quick notes
- If you're scraping a live website, always check the site's
robots.txtfile and terms of service to make sure you're allowed to scrape it - If the content loads dynamically (e.g., after clicking a button), you might need tools like
Seleniuminstead—but your example uses static HTML, soBeautifulSoupworks perfectly here
内容的提问来源于stack exchange,提问作者Aftab
相关产品推荐
相关产品推荐

