You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Beautiful Soup筛选class=high元素并提取Average Hourly Earnings m/m

实现步骤及代码

首先确保你已经安装了依赖库:

pip install beautifulsoup4 lxml

核心逻辑

你提供的HTML结构中,class="high"的节点和目标文本节点同属一个<tr>行,所以只需先定位high节点,再向上找到所属的行,再从行内提取事件标题文本即可,不会出现匹配错位的问题。

完整示例代码

from bs4 import BeautifulSoup

# 这里替换成你实际爬取到的HTML源码
html_content = '''
<tr class="calendar__row calendar_row calendar__row--grey calendar__row--no-grid nogrid" data-eventid="117390" data-ecobaseid="159" data-touchable="">
    <td class="calendar__cell calendar__impact impact calendar__impact calendar__impact--high">
        <div class="calendar__impact-icon calendar__impact-icon--screen">
            <span title="High Impact Expected" class="high"></span>
        </div>
        <div class="calendar__impact-icon calendar__impact-icon--print">
            <img src="https://resources.faireconomy.media/images/sprites/mm-impact-red.png" alt="" width="14" height="12">
        </div>
    </td>
    <td class="calendar__cell calendar__currency currency calendar__currency--right-of-impact" title="United States">
            US
    </td>   
    <td class="calendar__cell calendar__event event">
        <div>
            <span class="calendar__event-title">Average Hourly Earnings m/m</span>
        </div>
    </td>
</tr>
'''

# 初始化soup对象
soup = BeautifulSoup(html_content, 'lxml')

# 筛选所有class为high的节点
high_nodes = soup.find_all('span', class_='high')

for node in high_nodes:
    # 向上查找最近的tr父节点
    parent_tr = node.find_parent('tr')
    if not parent_tr:
        continue
    # 在当前tr内查找事件标题节点
    title_tag = parent_tr.find('span', class_='calendar__event-title')
    if title_tag:
        # 输出去除前后空白的文本内容
        print(title_tag.get_text(strip=True))

运行输出

Average Hourly Earnings m/m

内容的提问来源于stack exchange,提问作者Just Smile

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 04:15:02