You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无法从HTML Soup提取文本求助|附相关HTML代码片段

解决BeautifulSoup提取HTML文本的问题

没问题,我来帮你搞定这个提取文本的问题!先看你给出的HTML片段,咱们分情况来提取你需要的内容:

第一步:基础准备

首先确保你已经正确导入BeautifulSoup和解析器(比如html.parser),基础代码框架是这样的:

from bs4 import BeautifulSoup

# 把你的HTML片段赋值给变量
html_content = '''
<span class="tips yjSt" id="takane">その日はじめ(寄り付き)から現在までで、最も高かった値段</span></dt> </dl> </div> <div class="lineFi clearfix"> <dl class="tseDtl"><dd class="ymuiEditLink mar0"> <strong>189.1</strong><span class="date yjSt">(09:00)</span><span class="icoRealTime" title="リアルタイム"> </span></dd> <dt class="title">安値<a class="tips alignPos" data-ylk="slk:word;pos:4">
'''

# 初始化BeautifulSoup对象
soup = BeautifulSoup(html_content, 'html.parser')

第二步:提取目标文本

根据你需要的不同内容,用对应的选择器来提取:

  • 提取id为takane的span里的说明文本:
    这个元素有唯一的id,直接用find定位最靠谱:

    high_price_desc = soup.find('span', id='takane').get_text(strip=True)
    print(high_price_desc)
    # 输出:その日はじめ(寄り付き)から現在までで、最も高かった値段
    

    加strip=True可以去掉文本前后的空白字符,让结果更干净。

  • 提取strong标签里的数值189.1:
    直接定位strong标签即可:

    low_price_value = soup.find('strong').get_text(strip=True)
    print(low_price_value)
    # 输出:189.1
    
  • 提取dt标签里的「安値」文本:
    定位class为title的dt标签,然后提取文本:

    low_price_label = soup.find('dt', class_='title').get_text(strip=True)
    print(low_price_label)
    # 输出:安値
    

第三步:避免报错的小技巧

如果担心元素可能不存在(比如HTML结构变化),可以先判断元素是否存在再提取,避免AttributeError:

high_price_elem = soup.find('span', id='takane')
if high_price_elem:
    high_price_desc = high_price_elem.get_text(strip=True)
else:
    high_price_desc = "未找到目标元素"

内容的提问来源于stack exchange,提问作者Roman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 07:59:40