如何提取HTML中无锚标签的所有链接?含长API链接提取需求
解决方案
首先修正几个细节:你提供的LinkedIn URL拼写错误,正确地址是 https://www.linkedin.com/;另外目标链接里的somthing应该是something,下面的代码会按正确拼写处理。
核心思路
要提取无<a>标签的链接,需要分两步:
- 先收集所有
<a>标签的href属性(作为排除项) - 从HTML的其他标签属性、纯文本内容中匹配出所有URL,再排除掉已经在
<a>的href里的链接
同时使用更全面的正则表达式,匹配包含路径、参数的完整URL,解决之前只能匹配简单域名的问题。
完整代码
import requests import re from bs4 import BeautifulSoup from urllib.parse import urljoin # 修正后的目标URL url = 'https://www.linkedin.com/' # 目标API链接(修正拼写) target_api_link = 'https://api.something.com/v1/companies/' # 处理反爬,添加请求头 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } # 获取页面内容 response = requests.get(url, headers=headers) response.raise_for_status() # 检查请求是否成功 html_doc = response.text # 解析HTML soup = BeautifulSoup(html_doc, "html.parser") # 定义匹配完整URL的正则:支持http/https,包含域名、路径、参数 url_pattern = re.compile(r'https?://(?:[a-zA-Z0-9-]+\.)+[a-zA-Z]{2,}(?:/[^\s<>"]*)?') # 第一步:收集所有<a>标签的完整href链接(去重) a_hrefs = set() for a_tag in soup.find_all('a', href=True): href = a_tag['href'] # 转换为绝对URL,避免相对链接干扰 absolute_href = urljoin(url, href) if url_pattern.match(absolute_href): a_hrefs.add(absolute_href) # 第二步:收集所有非<a>标签的链接 non_a_links = set() # 遍历所有非<a>标签的属性,提取URL for tag in soup.find_all(): if tag.name == 'a': continue for attr_name, attr_value in tag.attrs.items(): if isinstance(attr_value, str): # 匹配属性值中的所有URL matches = url_pattern.findall(attr_value) for match in matches: # 转换为绝对URL并排除已在<a>中的链接 absolute_match = urljoin(url, match) if absolute_match not in a_hrefs: non_a_links.add(absolute_match) # 遍历所有非<a>包裹的文本节点,提取URL for text_node in soup.find_all(string=True): # 检查文本是否被<a>标签包裹 parent = text_node.parent is_in_a_tag = False while parent: if parent.name == 'a': is_in_a_tag = True break parent = parent.parent if not is_in_a_tag: matches = url_pattern.findall(text_node) for match in matches: absolute_match = urljoin(url, match) if absolute_match not in a_hrefs: non_a_links.add(absolute_match) # 提取目标API链接 found_target_links = [link for link in non_a_links if link == target_api_link] # 输出结果 print("=== 所有无<a>标签的链接 ===") for link in non_a_links: print(link) print("\n=== 找到的目标API链接 ===") if found_target_links: for link in found_target_links: print(link) else: print("未找到目标链接")
关键说明
- 正则表达式:
url_pattern可以匹配https://api.something.com/v1/companies/这类带多级路径的完整URL,解决了之前只能匹配简单域名的问题。 - 反爬处理:添加
User-Agent请求头,避免LinkedIn直接拒绝请求。 - 绝对URL转换:使用
urljoin将相对链接转换为绝对URL,保证链接的完整性。 - 去重处理:用集合存储链接,自动去除重复项。
- 精准过滤:通过检查文本节点的祖先标签,排除掉
<a>标签内部的文本链接,确保只提取真正无<a>包裹的链接。
内容的提问来源于stack exchange,提问作者zircon
相关产品推荐
相关产品推荐

