无法从HTML Soup提取文本求助|附相关HTML代码片段
解决BeautifulSoup提取HTML文本的问题
没问题,我来帮你搞定这个提取文本的问题!先看你给出的HTML片段,咱们分情况来提取你需要的内容:
第一步:基础准备
首先确保你已经正确导入BeautifulSoup和解析器(比如html.parser),基础代码框架是这样的:
from bs4 import BeautifulSoup # 把你的HTML片段赋值给变量 html_content = ''' <span class="tips yjSt" id="takane">その日はじめ(寄り付き)から現在までで、最も高かった値段</span></dt> </dl> </div> <div class="lineFi clearfix"> <dl class="tseDtl"><dd class="ymuiEditLink mar0"> <strong>189.1</strong><span class="date yjSt">(09:00)</span><span class="icoRealTime" title="リアルタイム"> </span></dd> <dt class="title">安値<a class="tips alignPos" data-ylk="slk:word;pos:4"> ''' # 初始化BeautifulSoup对象 soup = BeautifulSoup(html_content, 'html.parser')
第二步:提取目标文本
根据你需要的不同内容,用对应的选择器来提取:
提取id为
takane的span里的说明文本:
这个元素有唯一的id,直接用find定位最靠谱:high_price_desc = soup.find('span', id='takane').get_text(strip=True) print(high_price_desc) # 输出:その日はじめ(寄り付き)から現在までで、最も高かった値段加
strip=True可以去掉文本前后的空白字符,让结果更干净。提取strong标签里的数值
189.1:
直接定位strong标签即可:low_price_value = soup.find('strong').get_text(strip=True) print(low_price_value) # 输出:189.1提取dt标签里的「安値」文本:
定位class为title的dt标签,然后提取文本:low_price_label = soup.find('dt', class_='title').get_text(strip=True) print(low_price_label) # 输出:安値
第三步:避免报错的小技巧
如果担心元素可能不存在(比如HTML结构变化),可以先判断元素是否存在再提取,避免AttributeError:
high_price_elem = soup.find('span', id='takane') if high_price_elem: high_price_desc = high_price_elem.get_text(strip=True) else: high_price_desc = "未找到目标元素"
内容的提问来源于stack exchange,提问作者Roman
相关产品推荐
相关产品推荐

