如何用HTMLParser()提取指定层级下的目标price类标签内容?
解决HTMLParser提取最后一个指定类标签内容的问题
你的代码只能获取第一个price类div的内容,原因是找到第一个目标标签后就通过price_is_found标记停止了后续处理,同时没有正确管理标签的状态重置。以下是针对需求的修改方案:
核心思路
收集所有price类div的有效内容,最后取列表末尾的元素;同时正确处理标签的开始与结束状态,避免误捕获无效数据。
修改后的代码
from html.parser import HTMLParser class ParserLyku(HTMLParser): def __init__(self): super().__init__() self.is_price_field = False self._all_prices = [] def handle_starttag(self, tag, attrs): # 匹配div标签且class为price时,开启数据捕获标记 if tag == 'div': attrs_dict = dict(attrs) if attrs_dict.get('class') == 'price': self.is_price_field = True def handle_data(self, data): # 仅在price标签范围内时处理数据,过滤空白、::before等无效内容 if self.is_price_field: stripped_data = data.strip() if stripped_data and stripped_data != '::before': # 去除内容中的引号,保留纯数字 self._all_prices.append(stripped_data.strip('"')) def handle_endtag(self, tag): # 当price对应的div标签结束时,重置捕获标记 if tag == 'div' and self.is_price_field: self.is_price_field = False # 测试使用 html_content = """ <div class='price'>150</div> <div class='form-row'></div> <input type="hidden" value="15121" name="add-to-cart"> <div class='price'> ::before "1015" </div> """ parser = ParserLyku() parser.feed(html_content) # 取列表最后一个元素,即目标价格 if parser._all_prices: print(parser._all_prices[-1]) # 输出:1015
关键改动说明
- 移除了原代码中阻止后续处理的
price_is_found等判断,改为收集所有符合条件的price数据。 - 添加
handle_endtag方法,在div标签结束时重置is_price_field标记,避免后续非目标标签的数据被误捕获。 - 对捕获的数据做过滤处理,去掉空白字符、
::before和多余的引号,确保只保留有效内容。 - 最后从收集到的列表中取末尾元素,即为你需要的最后一个price标签内容。
如果目标price标签有特定的父级结构要求(比如必须在某个指定父标签下),还可以进一步扩展代码:在handle_starttag中跟踪当前标签层级,仅收集符合父级条件的price数据。
内容的提问来源于stack exchange,提问作者Beginner
相关产品推荐
相关产品推荐

