You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python网站颜色爬虫中CSS URL获取失败的问题求助

问题解决:抓取网站CSS颜色时的URL处理错误

你的代码在处理CSS文件URL时出错,核心原因是相对URL拼接逻辑错误:直接用urlunsplit给无协议的URL加http,会导致像/css/main.css这样的根相对路径被拼成http:///css/main.css(多了一个斜杠),触发InvalidURL错误;而对于完全相对的路径(比如css/style.css),这种拼接方式也无法正确关联到原网站域名。

修复方案

用urllib.parse.urljoin来拼接基础URL和CSS的相对URL,它能自动处理所有URL类型(相对路径、根路径、绝对路径),同时增加异常处理避免请求失败导致程序崩溃:

import re
import cssutils
import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin, urlsplit

url = 'https://www.endy.com/'

# 获取网站HTML
response = requests.get(url)
html = response.text

# 提取CSS文件URL(过滤无href的link标签)
soup = BeautifulSoup(html, 'html.parser')
css_urls = []
for link in soup.find_all('link', rel='stylesheet'):
    href = link.get('href')
    if href:
        css_urls.append(href)

color_dict = {}

# 处理每个CSS文件
for css_path in css_urls:
    # 用urljoin正确拼接基础URL和CSS路径
    css_url = urljoin(url, css_path)
    
    # 增加异常处理,避免请求失败中断程序
    try:
        css_response = requests.get(css_url)
        css_response.raise_for_status()  # 抛出HTTP错误
        css_text = css_response.text
        sheet = cssutils.parseString(css_text)

        # 提取颜色和选择器
        for rule in sheet:
            if rule.type == rule.STYLE_RULE:
                # 直接从style中提取颜色,不用拼接字符串再正则,更高效
                hex_colors = re.findall(r'#(?:[0-9a-fA-F]{3}){1,2}\b', rule.style.cssText)
                if hex_colors:
                    for color in hex_colors:
                        if color not in color_dict:
                            color_dict[color] = []
                        color_dict[color].append(rule.selectorText)
    except Exception as e:
        print(f"处理CSS {css_url} 失败: {str(e)}")
        continue

# 输出结果
for color, selectors in color_dict.items():
    print(f"颜色: {color}")
    print(f"选择器: {', '.join(selectors)}")
    print("------------------------------")

关键修改点

  • URL拼接:用urljoin(url, css_path)替代原逻辑,自动处理相对路径(如/assets/css/main.css)、相对路径(如css/style.css)和绝对路径(如https://cdn.example.com/style.css)。
  • 过滤无效标签:提取CSS URL时先判断href是否存在,避免空值导致错误。
  • 异常处理:用try-except包裹CSS请求和解析逻辑,单个CSS文件处理失败不会中断整个程序。
  • 优化颜色提取:直接从rule.style.cssText中提取颜色,不用拼接选择器和样式字符串,减少不必要的操作。

内容的提问来源于stack exchange,提问作者ANDREI MURESIAN

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 20:03:10