You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium Python爬取Realself图片及描述:当前方案是否可行?

Realself网站图片及描述爬取方案合理性分析

我是Selenium新手,想要保存Realself网站(目标页面:https://www.realself.com/photos/dermal-fillers#page=3&tags=&location=2038)每页的所有图片及对应描述。通过浏览器审查元素,得到目标内容的HTML结构如下:

<div class="fixed-img2">
 <div content="//fi.realself.com/340/2b7d179ededf559f87b06202a9187ee1/0/8/4/Injectable-Fillers-after-3292504-2757399.JPG" alt="Facial Rejuvenation with Injectable Fillers Facial rejuvenation was softly done over a course of 2 years to achieve a very natural and balanced" index="2" hide-icon="true" set-of-images="[&quot;//fi.realself.com/340/ff348ae5be9e09ede726574260fa386e/5/6/2/Injectable-Fillers-before-3292504-2757398.JPG&quot;]" regwall-override="true">
  <div class="Overlay--responsive">
    <div class="Overlay Overlay--explicitVideo Overlay--explicitBlock u-backgroundTransparent ng-hide" data-gtm="{&quot;event&quot;:&quot;safe-mode-block-click&quot;,&quot;safeMode&quot;:false}" ng-click="toggle()" ng-show="vm.blocked" role="button" tabindex="0" aria-hidden="true">
      <div ng-show="!vm.hideIcon" class="Overlay-explicitIcon ng-hide" style="" aria-hidden="true">
        <svg icon-id="nsfw-icon" class="Icon Icon--size56x56"><!----><use ng-if="$ctrl.iconURL" xlink:href="#rs-svg-nsfw-icon"></use><!----></svg>
      </div>
      <div class="Overlay-explicitTextBlock ng-hide" ng-show="vm.textmode" aria-hidden="true">
        <span class="Overlay-text Overlay-explicitTextBlack"> Sensitive content </span>
        <span class="Overlay-text Overlay-explicitTextBlack"> Click for real patient photos </span>
      </div>
    </div>
    
      <img src="//fi.realself.com/340/2b7d179ededf559f87b06202a9187ee1/0/8/4/Injectable-Fillers-after-3292504-2757399.JPG" class=" lazyloaded" data-src="//fi.realself.com/340/2b7d179ededf559f87b06202a9187ee1/0/8/4/Injectable-Fillers-after-3292504-2757399.JPG">
    
  </div>
</div>

目前我已经通过以下代码选中了所有class为fixed-img2的元素:

images = driver.find_elements(By.CLASS_NAME, "fixed-img2")

我的计划是遍历这些元素,提取子div的content属性获取图片链接,根据index属性区分术前/术后图片,并保存alt属性对应的图片描述。想请教这个方案是否合理?


方案整体合理性

这个思路完全可行,核心逻辑贴合页面结构,是针对当前HTML结构的直接解法,具体分析如下:

  • 图片链接提取:从子div的content属性取链接是准确的,注意要给链接补全https:前缀(原链接是相对协议的//开头),否则无法正常访问和下载。
  • 术前/术后区分:利用index属性区分是合理的,从示例HTML看index="2"对应术后图,而set-of-images里的是术前图,需要注意每个fixed-img2元素可能包含一组术前/术后图,遍历的时候别遗漏set-of-images里的链接(需要先转义字符串里的转义符,再解析成列表)。
  • 描述保存:alt属性的内容就是图片对应的描述,直接提取保存即可。

需要补充的注意事项

  • 元素等待:Selenium爬取时要确保元素完全加载,建议用WebDriverWait显式等待fixed-img2元素出现,避免因页面加载慢导致的空列表问题:
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

images = WebDriverWait(driver, 10).until(
    EC.presence_of_all_elements_located((By.CLASS_NAME, "fixed-img2"))
)
  • 反爬应对:Realself可能有反爬机制,建议添加随机延迟、更换User-Agent,避免短时间内频繁请求导致被封禁。
  • 图片下载:提取链接后可以用requests库下载图片,注意处理异常(比如链接失效、请求被拒),保存时可以用描述的关键词或者index命名,避免文件名重复。
  • 分页处理:要爬取多页的话,需要处理分页逻辑,比如找到分页按钮点击,或者修改URL里的page参数,注意判断是否到最后一页。

内容的提问来源于stack exchange,提问作者morteza eskandarian

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 11:05:11