You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:使用Scrapy爬取文章首图始终返回None的问题

问题:Scrapy爬取文章首图始终返回None无法获取链接

我用Python的Scrapy框架爬取文章首图(Feature Image),试了3-4种方法还是返回None,拿不到图片源链接,求帮忙排查。

我的代码:

class NewsSpider(scrapy.Spider):
    name = "cruisefever"

    def start_requests(self):
        url = input("Enter the article url: ")
        
        yield scrapy.Request(url, callback=self.parse_dir_contents)

    def parse_dir_contents(self, response):
        Feature_Image = response.xpath('//*[@id="td_uid_2_634abd2257025"]/div/div[1]/div/div[8]/div/p[1]/img/@data-src').extract()[0]
        #Feature_Image = response.xpath('//*[@id="td_uid_2_634abd2257025"]/div/div[1]/div/div[8]/div/p[1]/img/@data-img-url').extract()[0]
        #Feature_Image = response.xpath('//*[@id="td_uid_2_634abd2257025"]/div/div[1]/div/div[8]/div/p[1]/img/@src').extract()[0]
        #Feature_Image = [i.strip() for i in response.css('img[class*="alignnone size-full wp-image-39164 entered lazyloaded"] ::attr(src)').getall()][0]

        yield{
            'Feature_Image': Feature_Image,
        }

目标文章链接:https://cruisefever.net/carnival-cruise-lines-oldest-ship-sailing-final-cruise/


原因排查
  • 依赖动态ID:你用的XPath里包含td_uid_2_634abd2257025这种动态生成的ID,每次请求页面这个ID都会随机变化,导致XPath完全匹配不到目标元素。
  • CSS选择器依赖动态类名:wp-image-39164是当前文章图片的专属ID类,换其他文章就失效;而且lazyloaded是图片懒加载完成后才会添加的类,Scrapy默认获取的是页面初始HTML,此时这个类还没被加上。
  • 路径层级太固定:XPath里的div[8]这种固定层级写法很脆弱,一旦页面结构微调(比如新增/删除模块),路径就会失效。

解决方案

推荐两种稳定的获取方式:

方式1:通过og:image元数据获取(最通用)

页面的og:image元数据是专门为社交分享设置的首图标识,几乎不会随页面结构变动,通用性极强:

class NewsSpider(scrapy.Spider):
    name = "cruisefever"

    def start_requests(self):
        url = input("Enter the article url: ")
        yield scrapy.Request(url, callback=self.parse_dir_contents)

    def parse_dir_contents(self, response):
        # 提取og:image中的首图链接
        feature_image = response.xpath('//meta[@property="og:image"]/@content').get()
        
        yield {
            'Feature_Image': feature_image,
        }

方式2:通过文章内容区第一张图片获取

用页面稳定的类名定位文章内容容器,再提取里面的第一张图片链接:

class NewsSpider(scrapy.Spider):
    name = "cruisefever"

    def start_requests(self):
        url = input("Enter the article url: ")
        yield scrapy.Request(url, callback=self.parse_dir_contents)

    def parse_dir_contents(self, response):
        # 定位文章内容区,提取第一张图片的data-src(懒加载真实链接)
        feature_image = response.xpath('//div[@class="td-post-content tagdiv-type"]//p[1]/img/@data-src').get()
        # 如果data-src不存在, fallback到src
        if not feature_image:
            feature_image = response.xpath('//div[@class="td-post-content tagdiv-type"]//p[1]/img/@src').get()
        
        yield {
            'Feature_Image': feature_image,
        }

注意点

  • 用.get()替代.extract()[0]更安全,当匹配不到元素时会返回None,不会直接抛出索引错误。
  • 优先使用og:image,它是页面明确标记的首图,比从内容区找更准确。

内容的提问来源于stack exchange,提问作者Info Rewind

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 02:35:23