You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于BeautifulSoup的网页链接格式检测与关键词匹配技术需求

To spot links where the text starts with a digit sequence, followed right away by two distinct words (examples: "23 reasons to..." or "5 pictures to..."), we’ll use a regular expression to match this specific pattern.

Approach

  • Parse the HTML with BeautifulSoup to extract all link texts.
  • Use the regex pattern r'^\d+\s+\w+\s+\w+' to validate each link text:
    • ^: Ensures the match starts at the beginning of the string
    • \d+: Matches one or more digits
    • \s+: Matches one or more whitespace characters
    • \w+: Matches one or more word characters (letters, digits, underscores)

Code Implementation

from bs4 import BeautifulSoup
import re

def check_link_format(html_content):
    soup = BeautifulSoup(html_content, 'html5lib')
    links = soup.findAll('a')
    clean_links = [link.text.strip() for link in links if link.text.strip()]
    
    status = None
    pattern = re.compile(r'^\d+\s+\w+\s+\w+')
    
    for link_text in clean_links:
        if pattern.match(link_text):
            status = "ok"
            break  # Stop checking once we find a matching link
    
    return status

Building on your provided code snippet, we can optimize the keyword check to exit early as soon as a match is found, saving unnecessary iterations.

Code Implementation

from bs4 import BeautifulSoup

def check_link_keywords(html_content, keywords):
    soup = BeautifulSoup(html_content, 'html5lib')
    links = soup.findAll('a')
    clean_links = [link.text.strip() for link in links if link.text.strip()]
    
    status = None
    for link_text in clean_links:
        for keyword in keywords:
            # Optional: Make check case-insensitive for broader coverage
            if keyword.lower() in link_text.lower():
                status = "ok"
                return status  # Exit immediately once a match is found
    return status

Quick Notes

  • The case-insensitive check (lower()) helps catch matches like "Reasons" and "reasons" which are often the same in context.
  • Returning early cuts down on processing time, especially with large sets of links or keywords.

Combine Both Checks (Optional)

If you need to verify either the link format OR the presence of a keyword, you can merge the functions:

def check_links(html_content, keywords=None):
    # First check format
    format_status = check_link_format(html_content)
    if format_status == "ok":
        return "ok"
    
    # Then check keywords if provided
    if keywords:
        keyword_status = check_link_keywords(html_content, keywords)
        if keyword_status == "ok":
            return "ok"
    
    return None

内容的提问来源于stack exchange,提问作者Mathieu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 09:23:13