求助修正每日自动提取指定URL中心图片的Python代码
修正后的图片提取代码
原代码无法提取图片的核心问题是CSS选择器错误,#click-overlay并非图片标签,而是页面上的交互覆盖层,没有src属性。此外,部分网站会拦截无浏览器标识的请求,需要添加请求头模拟浏览器访问。
以下是修正后的完整代码:
from datetime import datetime, timedelta import requests from bs4 import BeautifulSoup import os import re def extract_image_url(page_url, img_selector): # 添加请求头模拟浏览器,避免被反爬拦截 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36' } response = requests.get(page_url, headers=headers) response.raise_for_status() # 检查下载是否成功 soup = BeautifulSoup(response.text, 'html.parser') # 使用正确的CSS选择器查找图片标签 img_tag = soup.select_one(img_selector) if img_tag and 'src' in img_tag.attrs: return img_tag['src'] else: # 备选方案:处理图片作为div背景图的情况 chart_div = soup.select_one('#chart-image') if chart_div and 'style' in chart_div.attrs: style_content = chart_div['style'] img_match = re.search(r'url\((.*?)\)', style_content) if img_match: return img_match.group(1).strip('\'"') raise ValueError("未找到图片或CSS选择器错误") def download_image(img_url, save_path): headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36' } # 完善相对路径转绝对路径逻辑 if not img_url.startswith('http'): img_url = f'https://weather.us{img_url}' if img_url.startswith('/') else f'https://weather.us/{img_url}' response = requests.get(img_url, headers=headers) response.raise_for_status() with open(save_path, 'wb') as f: f.write(response.content) print(f"图片已保存至:{save_path}") # 日期处理 today = datetime.now().strftime('%Y%m%d') dia_ahead = (datetime.now() + timedelta(days=16)).strftime('%Y%m%d') # URL构造(修正原URL中的时间格式错误) url_base = "https://weather.us/model-charts/standard/" url_today_00z_complete = f"{url_base}{today}00/brazil/accumulated-precipitation/{dia_ahead}-0000z.html" url_today_06z_complete = f"{url_base}{today}06/brazil/accumulated-precipitation/{dia_ahead}-0600z.html" print("\n今日00时完整URL:") print(url_today_00z_complete) print("\n今日06时完整URL:") print(url_today_06z_complete) urls = [url_today_00z_complete, url_today_06z_complete] # 替换为目标页面中中心图片的正确选择器 img_selector = "#chart-image img" # 提取并下载每个URL中的图片 for url in urls: try: img_url = extract_image_url(url, img_selector) save_path = os.path.join(os.getcwd(), os.path.basename(img_url)) download_image(img_url, save_path) except Exception as e: print(f"处理{url}时出错:{e}")
关键修正点:
- 修正CSS选择器:将无效的
#click-overlay替换为#chart-image img,匹配目标页面中的中心图片标签 - 添加请求头:模拟浏览器请求,绕过网站基础反爬拦截
- 完善路径处理:优化相对路径转绝对路径的逻辑,适配不同格式的图片URL
- 兼容背景图场景:增加正则提取背景图URL的备选方案,提升代码兼容性
- 修正URL格式:修复原代码中
06z对应的时间参数错误(将0060z改为0600z)
内容的提问来源于stack exchange,提问作者LeoLaboi
相关产品推荐
相关产品推荐

