如何提取网页链接并按是否存在nofollow标签分类标注
链接nofollow属性标注分类实现方法
nofollow是<a>标签rel属性的可选值,只要检测a标签的rel属性是否包含nofollow值即可完成标注。BeautifulSoup会自动将空格分隔的rel属性值拆分为列表,无需手动分割字符串处理。
修改要点
- 修正原代码的缩进问题,给BeautifulSoup指定
html.parser解析器,避免运行时的解析器警告 - 遍历a标签时同步提取
rel属性,无rel属性的标签默认判定为非nofollow - 将原存储纯链接字符串的列表改为存储结构化字典,同时记录链接地址、是否为nofollow两个属性,方便后续分类
- 补充全局nofollow检测:如果页面head存在
<meta name="robots" content="nofollow">标签,该页面下所有链接均属于nofollow链接
完整修改后代码
import httplib2 from bs4 import BeautifulSoup class Extractor(): def get_links(self, url): http = httplib2.Http() response, content = http.request(url) soup = BeautifulSoup(content, 'html.parser') links = [] # 检测页面全局nofollow规则 global_nofollow = False meta_robots = soup.find('meta', attrs={'name': 'robots'}) if meta_robots and meta_robots.get('content'): if 'nofollow' in meta_robots['content'].lower(): global_nofollow = True for link in soup.find_all('a', href=True): rel_vals = link.get('rel', []) is_nofollow = global_nofollow or ('nofollow' in rel_vals) links.append({ "url": link['href'], "is_nofollow": is_nofollow }) return links if __name__ == "__main__": url = '待爬取的目标网页地址' myextractor = Extractor() links = myextractor.get_links(url) # 分类输出 print("=== 普通follow链接 ===") for link in links: if not link['is_nofollow']: print(link['url']) print("\n=== nofollow链接 ===") for link in links: if link['is_nofollow']: print(link['url'])
使用说明
返回的links列表中每个元素都带is_nofollow布尔字段:
- 值为
False:普通可追踪链接,搜索引擎会传递对应权重 - 值为
True:nofollow链接,搜索引擎不会沿该链接传递权重
直接按该字段筛选即可完成两类链接的分类。
内容的提问来源于stack exchange,提问作者Joebevo
相关产品推荐
相关产品推荐

