You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过正则表达式忽略指定链接并精准统计文章内部链接数量?

Hey there! Let's get your link counting sorted out while excluding those category links. Your current regex approach won't work because regex doesn't use and like that—instead, we need to first capture the category links we want to ignore, then filter them out when counting.

First, your current code for grabbing category links isn't actually pulling the URLs. Let's fix that to get all the category URLs from the category div:

# Extract all category links from the div
category_div = soup.find('div', class_='category')
category_links = []
if category_div:
    category_links = [link.get('href') for link in category_div.find_all('a', href=True)]

We add a check for category_div to avoid errors if the div doesn't exist on some pages.

Instead of trying to cram exclusion logic into the regex, we can keep the regex to match your main domain links, then filter out any links that are in our category_links list. Here's how to adjust your link collection part:

# Keep the regex to match links from your main domain
pattern = re.compile(r"https://example.com/")

# Grab all main domain links first
all_main_links = [link.get('href') for link in body_text.find_all('a', href=pattern)]

# Filter out any links that are in the category_links list
filtered_links = [link for link in all_main_links if link not in category_links]

Full Updated Code

Putting it all together, your code will look like this:

import requests
from bs4 import BeautifulSoup
import re

links_bs4 = ['page1', 'page2']
data = []
all_links = []
pattern = re.compile(r"https://example.com/")

for item in links_bs4:
    page = requests.get(item)
    soup = BeautifulSoup(page.content, 'html.parser')
    
    # Get article title
    title = soup.find('title')
    title_text = title.string if title else "No Title"
    
    # Get article body
    body_text = soup.find('div', class_='article-body')
    if not body_text:
        data.append({'title': title_text, 'count links': 0})
        continue
    
    # Extract category links to ignore
    category_div = soup.find('div', class_='category')
    category_links = []
    if category_div:
        category_links = [link.get('href') for link in category_div.find_all('a', href=True)]
    
    # Collect and filter links
    all_main_links = [link.get('href') for link in body_text.find_all('a', href=pattern)]
    filtered_links = [link for link in all_main_links if link not in category_links]
    
    # Update data and all_links list
    data.append({'title': title_text, 'count links': len(filtered_links)})
    all_links.extend(filtered_links)

How This Works

  • We first pull all links from the article body that belong to your main domain (https://example.com/).
  • Then we exclude any links that match exactly with the ones found in the category div.
  • We added error handling for cases where the title, body, or category div might be missing (so your script doesn't crash unexpectedly).

For your example links (https://example.com/, https://example.com/category_1, https://example.com/category_2, https://example.com/page_1, https://example.com/page_2), this will only count page_1 and page_2—exactly what you want!

内容的提问来源于stack exchange,提问作者interferemadly

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 15:47:43