You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提取网页链接并按是否存在nofollow标签分类标注

链接nofollow属性标注分类实现方法

nofollow是<a>标签rel属性的可选值,只要检测a标签的rel属性是否包含nofollow值即可完成标注。BeautifulSoup会自动将空格分隔的rel属性值拆分为列表,无需手动分割字符串处理。

修改要点

  • 修正原代码的缩进问题,给BeautifulSoup指定html.parser解析器,避免运行时的解析器警告
  • 遍历a标签时同步提取rel属性,无rel属性的标签默认判定为非nofollow
  • 将原存储纯链接字符串的列表改为存储结构化字典,同时记录链接地址、是否为nofollow两个属性,方便后续分类
  • 补充全局nofollow检测:如果页面head存在<meta name="robots" content="nofollow">标签,该页面下所有链接均属于nofollow链接

完整修改后代码

import httplib2
from bs4 import BeautifulSoup

class Extractor():
    def get_links(self, url):
        http = httplib2.Http()
        response, content = http.request(url)
        soup = BeautifulSoup(content, 'html.parser')
        links = []

        # 检测页面全局nofollow规则
        global_nofollow = False
        meta_robots = soup.find('meta', attrs={'name': 'robots'})
        if meta_robots and meta_robots.get('content'):
            if 'nofollow' in meta_robots['content'].lower():
                global_nofollow = True

        for link in soup.find_all('a', href=True):
            rel_vals = link.get('rel', [])
            is_nofollow = global_nofollow or ('nofollow' in rel_vals)
            links.append({
                "url": link['href'],
                "is_nofollow": is_nofollow
            })
        return links


if __name__ == "__main__":
    url = '待爬取的目标网页地址'
    myextractor = Extractor()
    links = myextractor.get_links(url)

    # 分类输出
    print("=== 普通follow链接 ===")
    for link in links:
        if not link['is_nofollow']:
            print(link['url'])
    
    print("\n=== nofollow链接 ===")
    for link in links:
        if link['is_nofollow']:
            print(link['url'])

使用说明

返回的links列表中每个元素都带is_nofollow布尔字段:

  • 值为False:普通可追踪链接,搜索引擎会传递对应权重
  • 值为True:nofollow链接,搜索引擎不会沿该链接传递权重
    直接按该字段筛选即可完成两类链接的分类。

内容的提问来源于stack exchange,提问作者Joebevo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.30 09:21:34