求助:使用Scrapy爬取文章首图始终返回None的问题
问题:Scrapy爬取文章首图始终返回None无法获取链接
我用Python的Scrapy框架爬取文章首图(Feature Image),试了3-4种方法还是返回None,拿不到图片源链接,求帮忙排查。
我的代码:
class NewsSpider(scrapy.Spider): name = "cruisefever" def start_requests(self): url = input("Enter the article url: ") yield scrapy.Request(url, callback=self.parse_dir_contents) def parse_dir_contents(self, response): Feature_Image = response.xpath('//*[@id="td_uid_2_634abd2257025"]/div/div[1]/div/div[8]/div/p[1]/img/@data-src').extract()[0] #Feature_Image = response.xpath('//*[@id="td_uid_2_634abd2257025"]/div/div[1]/div/div[8]/div/p[1]/img/@data-img-url').extract()[0] #Feature_Image = response.xpath('//*[@id="td_uid_2_634abd2257025"]/div/div[1]/div/div[8]/div/p[1]/img/@src').extract()[0] #Feature_Image = [i.strip() for i in response.css('img[class*="alignnone size-full wp-image-39164 entered lazyloaded"] ::attr(src)').getall()][0] yield{ 'Feature_Image': Feature_Image, }
目标文章链接:https://cruisefever.net/carnival-cruise-lines-oldest-ship-sailing-final-cruise/
原因排查
- 依赖动态ID:你用的XPath里包含
td_uid_2_634abd2257025这种动态生成的ID,每次请求页面这个ID都会随机变化,导致XPath完全匹配不到目标元素。 - CSS选择器依赖动态类名:
wp-image-39164是当前文章图片的专属ID类,换其他文章就失效;而且lazyloaded是图片懒加载完成后才会添加的类,Scrapy默认获取的是页面初始HTML,此时这个类还没被加上。 - 路径层级太固定:XPath里的
div[8]这种固定层级写法很脆弱,一旦页面结构微调(比如新增/删除模块),路径就会失效。
解决方案
推荐两种稳定的获取方式:
方式1:通过og:image元数据获取(最通用)
页面的og:image元数据是专门为社交分享设置的首图标识,几乎不会随页面结构变动,通用性极强:
class NewsSpider(scrapy.Spider): name = "cruisefever" def start_requests(self): url = input("Enter the article url: ") yield scrapy.Request(url, callback=self.parse_dir_contents) def parse_dir_contents(self, response): # 提取og:image中的首图链接 feature_image = response.xpath('//meta[@property="og:image"]/@content').get() yield { 'Feature_Image': feature_image, }
方式2:通过文章内容区第一张图片获取
用页面稳定的类名定位文章内容容器,再提取里面的第一张图片链接:
class NewsSpider(scrapy.Spider): name = "cruisefever" def start_requests(self): url = input("Enter the article url: ") yield scrapy.Request(url, callback=self.parse_dir_contents) def parse_dir_contents(self, response): # 定位文章内容区,提取第一张图片的data-src(懒加载真实链接) feature_image = response.xpath('//div[@class="td-post-content tagdiv-type"]//p[1]/img/@data-src').get() # 如果data-src不存在, fallback到src if not feature_image: feature_image = response.xpath('//div[@class="td-post-content tagdiv-type"]//p[1]/img/@src').get() yield { 'Feature_Image': feature_image, }
注意点
- 用
.get()替代.extract()[0]更安全,当匹配不到元素时会返回None,不会直接抛出索引错误。 - 优先使用
og:image,它是页面明确标记的首图,比从内容区找更准确。
内容的提问来源于stack exchange,提问作者Info Rewind
相关产品推荐
相关产品推荐

