You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python和BeautifulSoup从HTML的data-bem中提取URL?

解决从data-bem属性解析URL的问题

问题排查与解决步骤:

  1. 确保Selenium获取完整渲染的页面
    页面元素可能是动态加载的,直接获取page_source可能会漏掉属性。先等待目标元素加载完成再提取源码:
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By

# 等待目标a标签加载完成(示例用class定位,可根据实际调整)
wait = WebDriverWait(driver, 15)
wait.until(EC.presence_of_element_located((By.CLASS_NAME, "OrganicTitle-Link")))

# 再获取完整页面源码
page_html = driver.page_source
  1. 正确解析转义后的data-bem内容
    HTML中的data-bem值是转义后的JSON格式(比如"是双引号的转义字符),需要先解码再解析:
from bs4 import BeautifulSoup
import json
import html

soup = BeautifulSoup(page_html, 'html.parser')
# 定位到目标a标签
target_link = soup.find('a', class_='OrganicTitle-Link')

if target_link:
    # 获取data-bem属性值
    bem_data = target_link.get('data-bem')
    if bem_data:
        # 解码HTML实体,还原JSON格式
        decoded_bem = html.unescape(bem_data)
        # 转为JSON对象
        bem_json = json.loads(decoded_bem)
        # 提取目标URL
        final_url = bem_json['click']['arguments']['url']
        print(final_url)  # 输出目标地址:https://allrival.com/parsing-saitov?yclid=1693735580326690815
    else:
        print("目标元素不存在data-bem属性")
else:
    print("未找到目标a标签")
  1. 验证元素定位准确性
    如果仍无法获取属性,检查你使用的定位方式(class、xpath等)是否精准匹配到目标元素,避免定位到没有data-bem的同类元素。

内容的提问来源于stack exchange,提问作者Rayqvaza

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 12:17:05