Python网站颜色爬虫中CSS URL获取失败的问题求助
问题解决:抓取网站CSS颜色时的URL处理错误
你的代码在处理CSS文件URL时出错,核心原因是相对URL拼接逻辑错误:直接用urlunsplit给无协议的URL加http,会导致像/css/main.css这样的根相对路径被拼成http:///css/main.css(多了一个斜杠),触发InvalidURL错误;而对于完全相对的路径(比如css/style.css),这种拼接方式也无法正确关联到原网站域名。
修复方案
用urllib.parse.urljoin来拼接基础URL和CSS的相对URL,它能自动处理所有URL类型(相对路径、根路径、绝对路径),同时增加异常处理避免请求失败导致程序崩溃:
import re import cssutils import requests from bs4 import BeautifulSoup from urllib.parse import urljoin, urlsplit url = 'https://www.endy.com/' # 获取网站HTML response = requests.get(url) html = response.text # 提取CSS文件URL(过滤无href的link标签) soup = BeautifulSoup(html, 'html.parser') css_urls = [] for link in soup.find_all('link', rel='stylesheet'): href = link.get('href') if href: css_urls.append(href) color_dict = {} # 处理每个CSS文件 for css_path in css_urls: # 用urljoin正确拼接基础URL和CSS路径 css_url = urljoin(url, css_path) # 增加异常处理,避免请求失败中断程序 try: css_response = requests.get(css_url) css_response.raise_for_status() # 抛出HTTP错误 css_text = css_response.text sheet = cssutils.parseString(css_text) # 提取颜色和选择器 for rule in sheet: if rule.type == rule.STYLE_RULE: # 直接从style中提取颜色,不用拼接字符串再正则,更高效 hex_colors = re.findall(r'#(?:[0-9a-fA-F]{3}){1,2}\b', rule.style.cssText) if hex_colors: for color in hex_colors: if color not in color_dict: color_dict[color] = [] color_dict[color].append(rule.selectorText) except Exception as e: print(f"处理CSS {css_url} 失败: {str(e)}") continue # 输出结果 for color, selectors in color_dict.items(): print(f"颜色: {color}") print(f"选择器: {', '.join(selectors)}") print("------------------------------")
关键修改点
- URL拼接:用
urljoin(url, css_path)替代原逻辑,自动处理相对路径(如/assets/css/main.css)、相对路径(如css/style.css)和绝对路径(如https://cdn.example.com/style.css)。 - 过滤无效标签:提取CSS URL时先判断
href是否存在,避免空值导致错误。 - 异常处理:用
try-except包裹CSS请求和解析逻辑,单个CSS文件处理失败不会中断整个程序。 - 优化颜色提取:直接从
rule.style.cssText中提取颜色,不用拼接选择器和样式字符串,减少不必要的操作。
内容的提问来源于stack exchange,提问作者ANDREI MURESIAN
相关产品推荐
相关产品推荐

