You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无法抓取Sunnxt剧集页面:请求响应失败及解析问题求助

Troubleshooting Scraping Issues with Sunnxt TV Detail Pages

Hey there! Let's work through this scraping problem with Sunnxt's TV detail pages together—you're trying to pull show names and air dates but hitting walls with your current Scrapy setup, right? Let's break down what might be going wrong and fix it step by step.

Common Roadblocks You're Likely Facing

  • Dynamic Content Rendering: Sunnxt, like most modern streaming platforms, uses JavaScript to load content after the initial page loads. Scrapy's default Request only grabs the static HTML, which doesn't include the show details you need.
  • Incomplete/Stale Request Headers: Your current headers include a User-Agent and csrf-token, but you might be missing critical headers (like Referer or Cookie) or using an expired csrf-token that's no longer valid.
  • Misaligned Selectors: If the content loads dynamically, your CSS/XPath selectors are targeting elements that don't exist in the initial static HTML.

Practical Solutions to Try

1. Use JavaScript Rendering with Scrapy Playwright

To get around dynamic content, integrate Scrapy with a tool that can fully render the page. Here's how to use Scrapy Playwright:

First, install the required package:

pip install scrapy-playwright

Update your settings.py to enable Playwright:

DOWNLOAD_HANDLERS = {
    "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
    "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
}

PLAYWRIGHT_LAUNCH_OPTIONS = {
    "headless": True,  # Set to False if you want to see the browser in action
    "args": ["--no-sandbox"],
}

Modify your spider to use Playwright requests:

import scrapy
from scrapy_playwright.page import PageCoroutine

class SunnxtSpider(scrapy.Spider):
    name = "sunnxt"
    start_urls = ["https://www.sunnxt.com/tv/detail/41600/"]

    def start_requests(self):
        for url in self.start_urls:
            yield scrapy.Request(
                url,
                meta={
                    "playwright": True,
                    "playwright_page_coroutines": [
                        # Wait for the show details container to load before parsing
                        PageCoroutine("wait_for_selector", "div.show-details-section"),
                    ],
                },
                headers={
                    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
                }
            )

    def parse(self, response):
        # Adjust selectors based on the actual rendered HTML (check browser Elements tab)
        show_name = response.css("h1.show-main-title::text").get().strip()
        air_date = response.css("span.show-air-date-text::text").get().strip()

        yield {
            "show_name": show_name,
            "air_date": air_date
        }

2. Bypass the Page—Fetch Data Directly from the API

Most streaming sites load content via hidden API calls. Open your browser's Network tab (look for XHR/fetch requests) while loading the TV detail page—you might find an API endpoint like https://api.sunnxt.com/v1/tv/detail/41600 that returns JSON with all the show data you need.

If you find this endpoint, you can skip scraping the page entirely:

import scrapy

class SunnxtAPISpider(scrapy.Spider):
    name = "sunnxt_api"
    # Replace with the actual API endpoint you find
    start_urls = ["https://api.sunnxt.com/v1/tv/detail/41600"]

    def parse(self, response):
        data = response.json()
        show_name = data["tv_show"]["title"]
        air_date = data["tv_show"]["first_air_date"]

        yield {
            "show_name": show_name,
            "air_date": air_date
        }

3. Fix CSRF-Token Handling (Don't Hardcode It!)

If the site requires a valid csrf-token, never hardcode it—tokens expire. Instead, grab it from the homepage first:

import scrapy

class SunnxtSpider(scrapy.Spider):
    name = "sunnxt"
    start_urls = ["https://www.sunnxt.com/"]

    def parse(self, response):
        # Extract fresh csrf-token from the homepage's meta tags
        csrf_token = response.css('meta[name="csrf-token"]::attr(content)').get()
        
        # Now request the TV detail page with the valid token
        yield scrapy.Request(
            "https://www.sunnxt.com/tv/detail/41600/",
            headers={
                "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36",
                "X-CSRF-Token": csrf_token
            },
            meta={"playwright": True},  # Keep this if dynamic content is still an issue
            callback=self.parse_detail
        )

    def parse_detail(self, response):
        show_name = response.css("h1.show-main-title::text").get().strip()
        air_date = response.css("span.show-air-date-text::text").get().strip()
        yield {"show_name": show_name, "air_date": air_date}

Quick Debugging Tips

  • Use scrapy shell to test requests and inspect the rendered response:
    scrapy shell "https://www.sunnxt.com/tv/detail/41600/" --playwright
    
    This lets you test CSS/XPath selectors interactively to make sure they target the right elements.
  • Check your browser's Elements tab after the page loads to find the exact selectors for show name and date.

内容的提问来源于stack exchange,提问作者Pradeep

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 03:37:51