You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Scrapy处理含JavaScript document动作的课程信息爬取?

How to Scrape JavaScript-Triggered Course Details in Scrapy

Got it, this is a super common issue when scraping sites that rely on JavaScript for navigation instead of traditional href links. Let's break down the practical approaches you can take to access those course details, just like you would with regular HTML:

1. Dig for the Real API Request (Easiest If Possible)

Most sites using JS to load content under the hood are making an AJAX/fetch request to a backend API. Here's how to track it down:

  • Open your browser's DevTools (F12), go to the Network tab, and filter for XHR/fetch requests.
  • Click the course number you want to scrape, and watch for a new request in the Network panel. This is likely the API call pulling the course details.
  • Copy the request URL, headers, and any parameters (like course ID), then use this directly in Scrapy instead of dealing with the JS.

Example Scrapy Code:

import scrapy
import json

class CourseSpider(scrapy.Spider):
    name = "course_spider"
    start_urls = ["https://example.com/course-list"]

    def parse(self, response):
        # Extract course IDs from the list page
        course_ids = response.css(".course-number::text").getall()
        for course_id in course_ids:
            # Use the API URL you found in DevTools
            api_url = f"https://example.com/api/courses/{course_id}"
            yield scrapy.Request(
                api_url,
                callback=self.parse_course_detail,
                meta={"course_id": course_id}
            )

    def parse_course_detail(self, response):
        # Parse the JSON response (most APIs return JSON)
        course_data = json.loads(response.text)
        yield {
            "course_id": response.meta["course_id"],
            "title": course_data["title"],
            "description": course_data["description"],
            "instructor": course_data["instructor"]
        }

2. Use Browser Automation (For Trickier Sites)

If the API is hidden, encrypted, or content is rendered directly via JS, use tools like Playwright or Selenium to let Scrapy control a real browser that executes the JS.

Setup with Scrapy-Playwright:

  1. Install the package: pip install scrapy-playwright
  2. Add these settings to your settings.py:
DOWNLOAD_HANDLERS = {
    "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
    "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
}
PLAYWRIGHT_LAUNCH_OPTIONS = {"headless": True}  # Set to False to see the browser

Example Spider Code:

import scrapy
from scrapy_playwright.page import PageCoroutine

class CourseSpider(scrapy.Spider):
    name = "course_spider"
    start_urls = ["https://example.com/course-list"]

    def start_requests(self):
        for url in self.start_urls:
            yield scrapy.Request(
                url,
                meta={
                    "playwright": True,
                    "playwright_page_coroutines": [
                        # Wait for the course list to load
                        PageCoroutine("wait_for_selector", ".course-item"),
                    ]
                }
            )

    def parse(self, response):
        # Iterate over course items and click each one
        for course_item in response.css(".course-item"):
            course_number = course_item.css(".course-number::text").get()
            yield scrapy.Request(
                response.url,
                callback=self.parse_course_detail,
                meta={
                    "playwright": True,
                    "playwright_page_coroutines": [
                        # Click the course number element
                        PageCoroutine("click", f".course-number:text-is('{course_number}')"),
                        # Wait for the detail modal/content to load
                        PageCoroutine("wait_for_selector", ".course-detail"),
                    ],
                    "course_number": course_number
                }
            )

    def parse_course_detail(self, response):
        yield {
            "course_number": response.meta["course_number"],
            "title": response.css(".course-detail h2::text").get(),
            "content": response.css(".course-detail .description::text").getall()
        }

3. Parse the JavaScript to Construct the Request

If the href looks like javascript:loadCourse('123') or javascript:document.getElementById('modal').loadData(456), extract the course ID from the JS code and build the corresponding URL manually.

Example Code to Extract JS Parameters:

import scrapy
import re

class CourseSpider(scrapy.Spider):
    name = "course_spider"
    start_urls = ["https://example.com/course-list"]

    def parse(self, response):
        for course_link in response.css("a[href^='javascript:']"):
            js_code = course_link.attrib["href"]
            # Use regex to extract the course ID from the JS function
            course_id_match = re.search(r"loadCourse\('(\d+)'\)", js_code)
            if course_id_match:
                course_id = course_id_match.group(1)
                # Construct the actual detail page URL (adjust based on what you find)
                detail_url = f"https://example.com/courses/detail?id={course_id}"
                yield scrapy.Request(detail_url, callback=self.parse_course_detail)

    def parse_course_detail(self, response):
        yield {
            "course_id": response.css(".course-id::text").get(),
            "title": response.css(".course-title::text").get()
        }

4. Check for Hidden DOM Content

Sometimes the site loads all course details into hidden HTML elements on the list page, and clicking just toggles visibility. Use Scrapy's selector to check for hidden divs/sections:

def parse(self, response):
    for course_item in response.css(".course-item"):
        # Extract hidden detail content directly from the list page
        hidden_detail = course_item.css(".hidden-course-detail::text").getall()
        yield {
            "course_number": course_item.css(".course-number::text").get(),
            "detail": hidden_detail
        }

内容的提问来源于stack exchange,提问作者Howard

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 08:24:03