You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

新手求助:如何从div标签内的span标签中提取数据?

Hey there! No worries at all—everyone starts with basic questions when diving into web scraping. Let's walk through exactly how to pull the data you need from that HTML snippet.

提取目标数据的分步指南

We'll use Python's BeautifulSoup library for this—it's super beginner-friendly and perfect for parsing static HTML like your example.

Step 1: Set up your environment

First, make sure you have the necessary libraries installed:

pip install beautifulsoup4 requests

(We'll use requests only if you're pulling HTML directly from a live website; if you already have the HTML snippet, you can skip the requests part.)

Step 2: Parse HTML and extract data

Here's a complete code example tailored to your HTML structure (I filled in the truncated old price part for clarity):

from bs4 import BeautifulSoup

# Your HTML snippet (with the truncated old price completed)
html_content = '''
<div class="price-container clearfix"> 
    <span class="sale-flag-percent">-40%</span> 
    <span class="price-box ri"> 
        <span class="price "><span data-currency-iso="PKR">Rs.</span> <span dir="ltr" data-price="5999">&nbsp;5,999</span> </span> 
        <span class="price -old "><span data-currency-iso="PKR">Rs.</span> <span dir="ltr" data-price="9999">&nbsp;9,999</span></span>
    </span>
</div>
'''

# Parse the HTML
soup = BeautifulSoup(html_content, 'html.parser')

# Locate the main price container div
price_container = soup.find('div', class_='price-container clearfix')

# Extract discount percentage
discount = price_container.find('span', class_='sale-flag-percent').text.strip()
print(f"Discount: {discount}")

# Extract current price details
current_price_section = price_container.find('span', class_='price ')
currency_symbol = current_price_section.find('span', attrs={'data-currency-iso': 'PKR'}).text.strip()
current_price_text = current_price_section.find('span', attrs={'data-price': True}).text.strip()
current_price_raw = current_price_section.find('span', attrs={'data-price': True})['data-price'] # Get pure numeric value
print(f"Current Price: {currency_symbol}{current_price_text} (Raw value: {current_price_raw})")

# Extract old price details
old_price_section = price_container.find('span', class_='price -old ')
old_currency_symbol = old_price_section.find('span', attrs={'data-currency-iso': 'PKR'}).text.strip()
old_price_text = old_price_section.find('span', attrs={'data-price': True}).text.strip()
old_price_raw = old_price_section.find('span', attrs={'data-price': True})['data-price']
print(f"Original Price: {old_currency_symbol}{old_price_text} (Raw value: {old_price_raw})")

Key explanations

  • soup.find(): Locates a single HTML element using attributes like class_ or custom attributes (like data-currency-iso)
  • .text.strip(): Pulls the visible text inside a tag and cleans up extra spaces/newlines
  • ['data-price']: Directly accesses the value of the data-price attribute, which gives you a clean numeric price (great for calculations later)

Quick notes

  • If you're scraping a live website, always check the site's robots.txt file and terms of service to make sure you're allowed to scrape it
  • If the content loads dynamically (e.g., after clicking a button), you might need tools like Selenium instead—but your example uses static HTML, so BeautifulSoup works perfectly here

内容的提问来源于stack exchange,提问作者Aftab

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:01:51