使用Selenium Python爬取Realself图片及描述:当前方案是否可行?
Realself网站图片及描述爬取方案合理性分析
我是Selenium新手,想要保存Realself网站(目标页面:https://www.realself.com/photos/dermal-fillers#page=3&tags=&location=2038)每页的所有图片及对应描述。通过浏览器审查元素,得到目标内容的HTML结构如下:
<div class="fixed-img2"> <div content="//fi.realself.com/340/2b7d179ededf559f87b06202a9187ee1/0/8/4/Injectable-Fillers-after-3292504-2757399.JPG" alt="Facial Rejuvenation with Injectable Fillers Facial rejuvenation was softly done over a course of 2 years to achieve a very natural and balanced" index="2" hide-icon="true" set-of-images="["//fi.realself.com/340/ff348ae5be9e09ede726574260fa386e/5/6/2/Injectable-Fillers-before-3292504-2757398.JPG"]" regwall-override="true"> <div class="Overlay--responsive"> <div class="Overlay Overlay--explicitVideo Overlay--explicitBlock u-backgroundTransparent ng-hide" data-gtm="{"event":"safe-mode-block-click","safeMode":false}" ng-click="toggle()" ng-show="vm.blocked" role="button" tabindex="0" aria-hidden="true"> <div ng-show="!vm.hideIcon" class="Overlay-explicitIcon ng-hide" style="" aria-hidden="true"> <svg icon-id="nsfw-icon" class="Icon Icon--size56x56"><!----><use ng-if="$ctrl.iconURL" xlink:href="#rs-svg-nsfw-icon"></use><!----></svg> </div> <div class="Overlay-explicitTextBlock ng-hide" ng-show="vm.textmode" aria-hidden="true"> <span class="Overlay-text Overlay-explicitTextBlack"> Sensitive content </span> <span class="Overlay-text Overlay-explicitTextBlack"> Click for real patient photos </span> </div> </div> <img src="//fi.realself.com/340/2b7d179ededf559f87b06202a9187ee1/0/8/4/Injectable-Fillers-after-3292504-2757399.JPG" class=" lazyloaded" data-src="//fi.realself.com/340/2b7d179ededf559f87b06202a9187ee1/0/8/4/Injectable-Fillers-after-3292504-2757399.JPG"> </div> </div>
目前我已经通过以下代码选中了所有class为fixed-img2的元素:
images = driver.find_elements(By.CLASS_NAME, "fixed-img2")
我的计划是遍历这些元素,提取子div的content属性获取图片链接,根据index属性区分术前/术后图片,并保存alt属性对应的图片描述。想请教这个方案是否合理?
方案整体合理性
这个思路完全可行,核心逻辑贴合页面结构,是针对当前HTML结构的直接解法,具体分析如下:
- 图片链接提取:从子div的
content属性取链接是准确的,注意要给链接补全https:前缀(原链接是相对协议的//开头),否则无法正常访问和下载。 - 术前/术后区分:利用
index属性区分是合理的,从示例HTML看index="2"对应术后图,而set-of-images里的是术前图,需要注意每个fixed-img2元素可能包含一组术前/术后图,遍历的时候别遗漏set-of-images里的链接(需要先转义字符串里的转义符,再解析成列表)。 - 描述保存:
alt属性的内容就是图片对应的描述,直接提取保存即可。
需要补充的注意事项
- 元素等待:Selenium爬取时要确保元素完全加载,建议用
WebDriverWait显式等待fixed-img2元素出现,避免因页面加载慢导致的空列表问题:
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC images = WebDriverWait(driver, 10).until( EC.presence_of_all_elements_located((By.CLASS_NAME, "fixed-img2")) )
- 反爬应对:Realself可能有反爬机制,建议添加随机延迟、更换User-Agent,避免短时间内频繁请求导致被封禁。
- 图片下载:提取链接后可以用
requests库下载图片,注意处理异常(比如链接失效、请求被拒),保存时可以用描述的关键词或者index命名,避免文件名重复。 - 分页处理:要爬取多页的话,需要处理分页逻辑,比如找到分页按钮点击,或者修改URL里的
page参数,注意判断是否到最后一页。
内容的提问来源于stack exchange,提问作者morteza eskandarian
相关产品推荐
相关产品推荐

