You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup4提取URN值及交易类型文本的技术问题

使用BeautifulSoup4提取目标内容方案

核心思路

针对你需要提取的三类内容,分别采用以下方式:

  • 以333开头的URN值:用正则匹配定位符合格式的文本或属性值
  • 交易类型文本:直接根据文本内容或标签特征精准查找

代码实现

from bs4 import BeautifulSoup
import re

# 替换为你的实际HTML内容
html_content = """
<div class="urn-box">
    <label>URN:</label>
    <span>333297706544052311</span>
</div>
<div class="transaction-actions">
    <button>Submit Renew</button>
    <h2>Labour Card Application</h2>
</div>
"""

# 解析HTML
soup = BeautifulSoup(html_content, 'html.parser')

# 提取URN值(匹配以333开头的纯数字字符串)
urn_regex = re.compile(r'^333\d+$')
urn = None
# 优先查找文本中的URN
for text_node in soup.find_all(text=urn_regex):
    urn = text_node.strip()
    break
# 如果URN在属性中(比如data-urn),启用下面的代码
# urn_element = soup.find(attrs={"data-urn": urn_regex})
# if urn_element:
#     urn = urn_element['data-urn']

# 提取交易类型相关文本
submit_renew_text = soup.find(text="Submit Renew").strip() if soup.find(text="Submit Renew") else None
labour_card_text = soup.find(text="Labour Card Application").strip() if soup.find(text="Labour Card Application") else None

# 输出结果
print(f"提取的URN: {urn}")
print(f"提取的Submit Renew文本: {submit_renew_text}")
print(f"提取的Labour Card Application文本: {labour_card_text}")

适配调整说明

  • 如果实际HTML中URN或交易文本的容器有特定class/id,可直接用soup.find('span', class_='urn-value')这类方式定位,更高效
  • 若目标文本存在多余空格,可在匹配时用re.compile(r'^\s*Submit Renew\s*$')忽略前后空白
  • 批量提取时,可改用find_all循环处理多个匹配项

内容的提问来源于stack exchange,提问作者Ranne Manuel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 08:56:05