You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用BeautifulSoup提取HTML中的文本?含示例处理需求

用BeautifulSoup提取HTML文本的实用方法

嘿,我来帮你搞定从这段HTML里提取内容的需求!针对你给出的HTML片段,我会一步步教你怎么用BeautifulSoup拿到像124这样的信息,还有其他可能需要的内容。

首先,先确保你已经安装了必要的库:

pip install beautifulsoup4 lxml

(用lxml作为解析器速度更快,当然也可以用Python自带的html.parser)

方法1:通过ID精准定位(最推荐)

因为你要找的数字124所在的<span>有唯一的id属性,直接通过id定位是最可靠的方式:

from bs4 import BeautifulSoup

# 你的HTML片段
html = '''<tr class="BgSilver" style="border-color:Gray;border-width:1px;border-style:Solid;"> <td align="right" style="width:75px;" valign="top"> <span id="ctl00_cph1_grdAwardSearch_ctl26_lblRowNum" style="display:inline-block;width:50px;">124</span> </td> <td align="left" valign="top"> <span id="ctl00_cph1_grdAwardSearch_ctl26_lblAwardBasicNumber" style="display:inline-block;width:150px;"><a href="https://dibbs...">奖项编号</a></span> </td></tr>'''

# 初始化BeautifulSoup对象
soup = BeautifulSoup(html, 'lxml')

# 通过id提取目标span的文本
row_num = soup.find('span', id='ctl00_cph1_grdAwardSearch_ctl26_lblRowNum').get_text(strip=True)
print(row_num)  # 输出:124

这里的get_text(strip=True)会自动去掉文本前后的空格和换行,非常实用。

方法2:通过层级结构定位(如果id会动态变化)

如果这个表格行的id是动态生成的(比如ctl26会变),那可以通过父元素的class来定位:

# 先找到class为BgSilver的tr标签
target_tr = soup.find('tr', class_='BgSilver')
# 然后找到第一个td里的span文本
row_num = target_tr.find('td').find('span').get_text(strip=True)
print(row_num)  # 输出:124

或者更简洁的链式写法:

row_num = soup.find('tr', class_='BgSilver').td.span.get_text(strip=True)

顺便提取其他内容(比如奖项编号链接)

如果你还需要提取那个<a>标签里的文本或者链接,也很简单:

# 提取链接文本
award_num_text = soup.find('span', id='ctl00_cph1_grdAwardSearch_ctl26_lblAwardBasicNumber').a.get_text(strip=True)
# 提取链接地址
award_link = soup.find('span', id='ctl00_cph1_grdAwardSearch_ctl26_lblAwardBasicNumber').a['href']

这些方法基本上能覆盖你从这段HTML里提取内容的需求啦,要是有其他变体情况,随时调整定位逻辑就行~

内容的提问来源于stack exchange,提问作者jone2

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:49:56