You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提取HTML中无锚标签的所有链接?含长API链接提取需求

解决方案

首先修正几个细节:你提供的LinkedIn URL拼写错误,正确地址是 https://www.linkedin.com/;另外目标链接里的somthing应该是something,下面的代码会按正确拼写处理。

核心思路

要提取无<a>标签的链接,需要分两步:

  1. 先收集所有<a>标签的href属性(作为排除项)
  2. 从HTML的其他标签属性、纯文本内容中匹配出所有URL,再排除掉已经在<a>的href里的链接

同时使用更全面的正则表达式,匹配包含路径、参数的完整URL,解决之前只能匹配简单域名的问题。

完整代码

import requests
import re
from bs4 import BeautifulSoup
from urllib.parse import urljoin

# 修正后的目标URL
url = 'https://www.linkedin.com/'
# 目标API链接(修正拼写)
target_api_link = 'https://api.something.com/v1/companies/'

# 处理反爬,添加请求头
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
}

# 获取页面内容
response = requests.get(url, headers=headers)
response.raise_for_status()  # 检查请求是否成功
html_doc = response.text

# 解析HTML
soup = BeautifulSoup(html_doc, "html.parser")

# 定义匹配完整URL的正则:支持http/https,包含域名、路径、参数
url_pattern = re.compile(r'https?://(?:[a-zA-Z0-9-]+\.)+[a-zA-Z]{2,}(?:/[^\s<>"]*)?')

# 第一步:收集所有<a>标签的完整href链接(去重)
a_hrefs = set()
for a_tag in soup.find_all('a', href=True):
    href = a_tag['href']
    # 转换为绝对URL,避免相对链接干扰
    absolute_href = urljoin(url, href)
    if url_pattern.match(absolute_href):
        a_hrefs.add(absolute_href)

# 第二步:收集所有非<a>标签的链接
non_a_links = set()

# 遍历所有非<a>标签的属性,提取URL
for tag in soup.find_all():
    if tag.name == 'a':
        continue
    for attr_name, attr_value in tag.attrs.items():
        if isinstance(attr_value, str):
            # 匹配属性值中的所有URL
            matches = url_pattern.findall(attr_value)
            for match in matches:
                # 转换为绝对URL并排除已在<a>中的链接
                absolute_match = urljoin(url, match)
                if absolute_match not in a_hrefs:
                    non_a_links.add(absolute_match)

# 遍历所有非<a>包裹的文本节点,提取URL
for text_node in soup.find_all(string=True):
    # 检查文本是否被<a>标签包裹
    parent = text_node.parent
    is_in_a_tag = False
    while parent:
        if parent.name == 'a':
            is_in_a_tag = True
            break
        parent = parent.parent
    if not is_in_a_tag:
        matches = url_pattern.findall(text_node)
        for match in matches:
            absolute_match = urljoin(url, match)
            if absolute_match not in a_hrefs:
                non_a_links.add(absolute_match)

# 提取目标API链接
found_target_links = [link for link in non_a_links if link == target_api_link]

# 输出结果
print("=== 所有无<a>标签的链接 ===")
for link in non_a_links:
    print(link)

print("\n=== 找到的目标API链接 ===")
if found_target_links:
    for link in found_target_links:
        print(link)
else:
    print("未找到目标链接")

关键说明

  1. 正则表达式:url_pattern可以匹配https://api.something.com/v1/companies/这类带多级路径的完整URL,解决了之前只能匹配简单域名的问题。
  2. 反爬处理:添加User-Agent请求头,避免LinkedIn直接拒绝请求。
  3. 绝对URL转换:使用urljoin将相对链接转换为绝对URL,保证链接的完整性。
  4. 去重处理:用集合存储链接,自动去除重复项。
  5. 精准过滤:通过检查文本节点的祖先标签,排除掉<a>标签内部的文本链接,确保只提取真正无<a>包裹的链接。

内容的提问来源于stack exchange,提问作者zircon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 01:57:56