You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup提取链接时,如何仅获取纯链接地址去除多余属性?

解决方案

你可以通过提取每个<a>标签的href属性值来获取纯链接地址,以下是两种实用修改方式:

方式一:直接提取属性值(简洁高效)

修改后的完整代码:

import requests
from bs4 import BeautifulSoup
import re

url = 'https://link.com/'

html = requests.get(url)

soup = BeautifulSoup(html.content, 'html.parser')

# 获取符合条件的a标签列表
link_tags = soup.findAll('a', href=re.compile('https://specificlink/'))
# 遍历列表提取每个标签的href属性
pure_links = [tag['href'] for tag in link_tags]

print(pure_links)

方式二:用get()方法避免报错(更健壮)

如果存在部分<a>标签缺失href属性的情况,使用get()方法可以避免抛出异常,缺失时返回None:

import requests
from bs4 import BeautifulSoup
import re

url = 'https://link.com/'

html = requests.get(url)

soup = BeautifulSoup(html.content, 'html.parser')

link_tags = soup.findAll('a', href=re.compile('https://specificlink/'))
pure_links = [tag.get('href') for tag in link_tags]

print(pure_links)

执行后输出示例:

['https://specificlink']

内容的提问来源于stack exchange,提问作者Krasav4ic

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 03:40:16