使用BeautifulSoup4提取URN值及交易类型文本的技术问题
使用BeautifulSoup4提取目标内容方案
核心思路
针对你需要提取的三类内容,分别采用以下方式:
- 以333开头的URN值:用正则匹配定位符合格式的文本或属性值
- 交易类型文本:直接根据文本内容或标签特征精准查找
代码实现
from bs4 import BeautifulSoup import re # 替换为你的实际HTML内容 html_content = """ <div class="urn-box"> <label>URN:</label> <span>333297706544052311</span> </div> <div class="transaction-actions"> <button>Submit Renew</button> <h2>Labour Card Application</h2> </div> """ # 解析HTML soup = BeautifulSoup(html_content, 'html.parser') # 提取URN值(匹配以333开头的纯数字字符串) urn_regex = re.compile(r'^333\d+$') urn = None # 优先查找文本中的URN for text_node in soup.find_all(text=urn_regex): urn = text_node.strip() break # 如果URN在属性中(比如data-urn),启用下面的代码 # urn_element = soup.find(attrs={"data-urn": urn_regex}) # if urn_element: # urn = urn_element['data-urn'] # 提取交易类型相关文本 submit_renew_text = soup.find(text="Submit Renew").strip() if soup.find(text="Submit Renew") else None labour_card_text = soup.find(text="Labour Card Application").strip() if soup.find(text="Labour Card Application") else None # 输出结果 print(f"提取的URN: {urn}") print(f"提取的Submit Renew文本: {submit_renew_text}") print(f"提取的Labour Card Application文本: {labour_card_text}")
适配调整说明
- 如果实际HTML中URN或交易文本的容器有特定class/id,可直接用
soup.find('span', class_='urn-value')这类方式定位,更高效 - 若目标文本存在多余空格,可在匹配时用
re.compile(r'^\s*Submit Renew\s*$')忽略前后空白 - 批量提取时,可改用
find_all循环处理多个匹配项
内容的提问来源于stack exchange,提问作者Ranne Manuel
相关产品推荐
相关产品推荐

