如何用Beautiful Soup定位HTML中data-bind标签并提取指定文本
用Beautiful Soup提取带data-bind属性元素的文本内容
你提供的目标HTML代码如下:
<span class="price hepsiburada price-new-old" itemprop="price" data-bind="css: {'merchant': !isHepsiburadaProduct(), 'hepsiburada': isHepsiburadaProduct(), 'price-new-old': product().currentListing.pricing.listingPriceList.length === 1, 'price-new': product().currentListing.pricing.listingPriceList[product().currentListing.pricing.listingPriceList.length - 1].discountRate > 0}" id="offering-price" content="2179.00"> <span data-bind="markupText:'currentPriceBeforePoint'">2.179</span>,<span data-bind="markupText:'currentPriceAfterPoint'">00</span> <span class="turkishLira" itemprop="priceCurrency" content="TRY">TL</span> <span class="hidden" style="font-weight: normal; font-size: 16px; color: #646464; " data-bind="css: {hidden: unitPriceFormatted() == ''}, html:unitPriceFormatted()"></span> </span>
要提取文本2.179,可以通过属性选择器直接定位带有指定data-bind属性的<span>元素,Beautiful Soup完全支持这种方式,以下是两种可行方案:
方案1:使用CSS选择器
直接通过[data-bind="属性值"]的语法匹配目标元素,代码示例:
from bs4 import BeautifulSoup # 假设html是你获取到的HTML内容 soup = BeautifulSoup(html, 'html.parser') # 定位data-bind值为markupText:'currentPriceBeforePoint'的span target_span = soup.select_one('[data-bind="markupText:\'currentPriceBeforePoint\'"]') # 提取文本内容 price_before_point = target_span.get_text(strip=True) print(price_before_point) # 输出:2.179
注意:属性值包含单引号时,外层用双引号包裹或对单引号转义,避免语法冲突。
方案2:使用find方法的attrs参数
通过find()方法的attrs字典参数指定要匹配的data-bind属性,代码示例:
from bs4 import BeautifulSoup soup = BeautifulSoup(html, 'html.parser') # 用attrs字典匹配data-bind属性 target_span = soup.find('span', attrs={'data-bind': "markupText:'currentPriceBeforePoint'"}) price_before_point = target_span.get_text(strip=True) print(price_before_point) # 输出:2.179
如果页面存在多个相同data-bind值的元素,可替换为select()或find_all()遍历筛选。
内容的提问来源于stack exchange,提问作者illans
相关产品推荐
相关产品推荐

