You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python结合Beautiful Soup获取按钮的data-src数据?

使用Beautiful Soup提取按钮的data-src属性值

要批量收集网页中所有按钮的data-src属性值,用Beautiful Soup可以这么操作:

步骤说明

  • 导入BeautifulSoup库,若需爬取远程网页,额外导入requests
  • 解析网页HTML内容
  • 定位所有目标按钮元素
  • 逐一提取data-src属性值并收集

代码示例

from bs4 import BeautifulSoup
# 爬远程网页时需导入requests
# import requests

# 场景1:使用本地HTML字符串(对应你提供的示例代码)
html_content = '''
<button aria-label="Play" class="history" data-src="webpage.com" data-type="audio/mp3" title="Play" type="button"></button>
<button aria-label="Play" class="history" data-src="another-audio.com" data-type="audio/mp3" title="Play" type="button"></button>
'''

# 场景2:爬取远程网页的HTML内容
# response = requests.get("目标网页URL")
# html_content = response.text

# 解析HTML,默认用html.parser即可,也可使用lxml(需额外安装)
soup = BeautifulSoup(html_content, 'html.parser')

# 定位目标按钮:两种方式二选一
# 方式1:匹配所有class为history的button元素(和你示例中的按钮精准匹配)
target_buttons = soup.find_all('button', class_='history')

# 方式2:更通用,匹配所有带有data-src属性的button元素
# target_buttons = soup.find_all('button', attrs={'data-src': True})

# 收集所有有效的data-src值
collected_links = []
for btn in target_buttons:
    # 用get方法提取属性,避免元素无该属性时抛出异常
    src_link = btn.get('data-src')
    if src_link:  # 确保值不为空再添加到列表
        collected_links.append(src_link)

# 输出结果
print(collected_links)
# 输出示例:['webpage.com', 'another-audio.com']

关键提示

  • 用class_而非class作为参数,因为class是Python关键字,直接使用会报错
  • btn.get('data-src')比btn['data-src']更安全,当元素没有data-src属性时,前者返回None,后者会触发KeyError
  • 若网页内容是JS动态渲染的,Beautiful Soup无法直接获取,需用Selenium或Playwright等工具先渲染页面再提取

内容的提问来源于stack exchange,提问作者goodkidbadmAArky

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 21:01:00