You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python Scrapy爬取Travel + Leisure文章:解决媒体URL与内容爬取问题

Travel + Leisure网站Scrapy爬虫问题解决方案

问题描述

我希望使用Python的Scrapy框架爬取Travel + Leisure指定网站的文章,需实现以下需求:

  • 获取文章的特色视频URL或特色图片URL;
  • 按文章原有顺序提取所有标题、段落内容;
  • 排除主文章div(class="loc article-content")内的非文章相关文本。

目前无法获取特色媒体URL,且提取文章文本时遇到问题,现有爬虫代码如下:

from urllib.parse import urljoin
import scrapy
from scrapy.crawler import CrawlerProcess
from scrapy.utils.project import get_project_settings
from datetime import datetime
import pandas as pd


class NewsSpider(scrapy.Spider):
    name = "travelandleisure"

    def start_requests(self):
        url = input("Enter the article url: ")
        
        yield scrapy.Request(url, callback=self.parse_dir_contents)

    def parse_dir_contents(self, response):
        try:
            Author = ', '.join(set([x.strip() for x in response.xpath('//a[@class="mntl-attribution__item-name"]/text()').extract()]))
        except IndexError:
            Author = "NULL"
        
        url = response.url
        
        try:
            Category = response.xpath('//*[@id="mntl-text-link_1-0"]/span/text()').get()
        except IndexError:
            Category = "NULL"

        Headlines = response.xpath('//*[@id="article-heading_1-0"]/text()').get().replace("\n","")

        
        
        Source = response.xpath('//*[@id="mntl-text-block_1-0"]/text()').get().replace("\n", "")
        
        Published_Date = response.css('div.mntl-attribution__item-date::text').get().split("on ")[1].replace(",","")#Updated on June 8, 2022
        Published_Date = datetime.strptime(Published_Date, "%B %d %Y").date()
        
        #================Waiting for Stack answer====================
        
        Feature_Image = "NULL" #Please tell the code for feature Image or Video
        
        Content =  "NULL" #Please tell the code for Scrape the all text, paragraph and heading but in sorting as in the article but not include the text that not belong to article but wrapped in this article div " <div class="loc article-content"> "

        yield{
            'Category':Category,
            'Headlines':Headlines,
            'Author': Author,
            'Source': Source,
            'Publication Date': Published_Date,
            'Feature_Image': Feature_Image,
            'Skift Take': skift_take,
            'Article Content': Content
        }
            # =============== Data Store +++++++++++++++++++++
        Data = [[Category,Headlines,Author,Source,Published_Date,Feature_Image,Content,url]]

        cols = ['Category','Headlines','Author','Source','Published_Date','Feature_Image','Content','URL']
        try:
            opened_df = pd.read_csv('C:/Users/Public/pagedata.csv')
            opened_df = pd.concat([opened_df,pd.DataFrame(Data, columns = cols)])
        except:
            opened_df = pd.DataFrame(Data, columns = cols)

        opened_df.to_csv('C:/Users/Public/pagedata.csv', index= False) 

if __name__ == '__main__':
    settings = get_project_settings()
    process = CrawlerProcess(settings)
    process.crawl(NewsSpider)
    
    process.start()

目标文章URL:https://www.travelandleisure.com/travel-news/where-can-americans-travel-right-now-a-country-by-country-guide


解决方案

1. 特色媒体URL获取逻辑

Travel + Leisure的特色媒体优先放在视频容器或主图片容器中,代码逻辑如下:

  • 先检查是否存在视频容器(div.mntl-video-player),提取其data-src属性作为视频URL;
  • 若没有视频,检查主图片容器(div.mntl-primary-image-container),优先从img标签的data-srcset中取最高分辨率图片地址,备用取src属性。

2. 文章内容提取逻辑

限定在主文章区域div.loc.article-content内,只提取标题(h2、h3)和段落(p)元素,同时排除广告、相关推荐等非内容模块(如带mntl-ad-unit、mntl-related-content类的元素),按原文顺序拼接内容,给标题加标识区分层级。

修改后的完整代码

from urllib.parse import urljoin
import scrapy
from scrapy.crawler import CrawlerProcess
from scrapy.utils.project import get_project_settings
from datetime import datetime
import pandas as pd


class NewsSpider(scrapy.Spider):
    name = "travelandleisure"

    def start_requests(self):
        url = input("Enter the article url: ")
        yield scrapy.Request(url, callback=self.parse_dir_contents)

    def parse_dir_contents(self, response):
        try:
            Author = ', '.join(set([x.strip() for x in response.xpath('//a[@class="mntl-attribution__item-name"]/text()').extract()]))
        except IndexError:
            Author = "NULL"
        
        url = response.url
        
        try:
            Category = response.xpath('//*[@id="mntl-text-link_1-0"]/span/text()').get()
        except IndexError:
            Category = "NULL"

        Headlines = response.xpath('//*[@id="article-heading_1-0"]/text()').get().replace("\n","")

        # 优化Source的异常处理
        try:
            Source = response.xpath('//*[@id="mntl-text-block_1-0"]/text()').get().replace("\n", "")
        except AttributeError:
            Source = "NULL"
        
        # 优化发布日期的异常处理
        try:
            Published_Date = response.css('div.mntl-attribution__item-date::text').get().split("on ")[1].replace(",","")
            Published_Date = datetime.strptime(Published_Date, "%B %d %Y").date()
        except (IndexError, AttributeError):
            Published_Date = "NULL"
        
        # 获取特色媒体URL(优先视频,再图片)
        Feature_Media = "NULL"
        # 提取视频URL
        video_src = response.css('div.mntl-video-player[data-src]::attr(data-src)').get()
        if video_src:
            Feature_Media = urljoin(response.url, video_src)
        else:
            # 提取图片URL,优先取高清的srcset地址
            img_srcset = response.css('div.mntl-primary-image-container img::attr(data-srcset)').get()
            if img_srcset:
                # 取srcset中最后一个(通常是最高分辨率)地址
                Feature_Media = img_srcset.split(',')[-1].split()[0].strip()
            else:
                # 备用取src属性
                img_src = response.css('div.mntl-primary-image-container img::attr(src)').get()
                if img_src:
                    Feature_Media = urljoin(response.url, img_src)

        # 提取文章内容:按顺序获取标题和段落,排除非内容元素
        content_parts = []
        article_content = response.css('div.loc.article-content')
        # 遍历文章内的标题和段落元素
        for element in article_content.xpath('.//*[self::h2 or self::h3 or self::p]'):
            # 跳过广告、相关推荐等非内容模块的元素
            if element.xpath('./ancestor::div[contains(@class, "mntl-ad-unit") or contains(@class, "mntl-related-content")]'):
                continue
            text = element.xpath('string(.)').get().strip()
            if text:
                # 给标题加标识区分
                if element.xpath('self::h2') or element.xpath('self::h3'):
                    content_parts.append(f"【标题】{text}")
                else:
                    content_parts.append(text)
        Content = '\n\n'.join(content_parts) if content_parts else "NULL"

        # 移除原代码中未定义的skift_take字段
        yield{
            'Category':Category,
            'Headlines':Headlines,
            'Author': Author,
            'Source': Source,
            'Publication Date': Published_Date,
            'Feature_Media': Feature_Media,
            'Article Content': Content
        }
        
        # 数据存储逻辑
        Data = [[Category,Headlines,Author,Source,Published_Date,Feature_Media,Content,url]]
        cols = ['Category','Headlines','Author','Source','Published_Date','Feature_Media','Content','URL']
        
        try:
            opened_df = pd.read_csv('C:/Users/Public/pagedata.csv')
            opened_df = pd.concat([opened_df,pd.DataFrame(Data, columns = cols)])
        except:
            opened_df = pd.DataFrame(Data, columns = cols)
        
        opened_df.to_csv('C:/Users/Public/pagedata.csv', index= False) 

if __name__ == '__main__':
    settings = get_project_settings()
    process = CrawlerProcess(settings)
    process.crawl(NewsSpider)
    
    process.start()

关键优化点

  • 新增特色媒体(视频/图片)的提取逻辑,覆盖两种媒体类型;
  • 优化文章内容提取,过滤非内容元素,保留原文顺序和层级;
  • 增加多处异常处理,避免因元素不存在导致爬虫崩溃;
  • 移除原代码中未定义的skift_take字段,修复语法错误。

内容的提问来源于stack exchange,提问作者Info Rewind

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 14:25:19