You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无法通过CSS/XPath提取网页CSRF Token,请求技术协助

Troubleshooting CSRF Token Extraction from Sunnxt Movie Page

Hey there! The selectors you’re using (meta[name="csrf-token"]::attr(content) for CSS and //meta[@name="csrf-token"]/@content for XPath) are actually perfectly correct—so the issue is almost certainly that the page loads this meta tag dynamically using JavaScript. Static HTML-fetching tools (like Python’s requests library) only grab the initial server response, which doesn’t include elements injected after the page loads via client-side code.

Here’s a breakdown of the problem and fixes:

Why Your Current Approach Isn’t Working

If you right-click the page and select "View Page Source", you probably won’t see the csrf-token meta tag listed there. But if you use your browser’s dev tools (F12) to inspect the DOM, you will see it. That’s a clear sign the token is added dynamically by JavaScript after the initial HTML loads. Tools that don’t execute JS can’t access elements added this way.

Solutions Using JavaScript-Rendering Tools

You’ll need to use a tool that simulates a real browser (and runs JS) to access the fully rendered DOM. Here are a few popular options with code examples:

Solution 1: Python + Selenium

Selenium controls a real browser, so it renders the page exactly like a human would.

First, make sure you have Selenium and ChromeDriver installed (or use another browser like Firefox):

pip install selenium
# Download ChromeDriver from https://sites.google.com/chromium.org/driver/ and add it to your PATH

Then the code:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# Initialize the browser
driver = webdriver.Chrome()

try:
    driver.get("https://www.sunnxt.com/movie/inside/")
    # Wait up to 10 seconds for the meta tag to appear in the DOM
    csrf_meta = WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CSS_SELECTOR, 'meta[name="csrf-token"]'))
    )
    csrf_token = csrf_meta.get_attribute("content")
    print("Extracted CSRF Token:", csrf_token)
finally:
    # Always close the browser when done
    driver.quit()

Solution 2: Python + Playwright

Playwright is a modern, fast alternative to Selenium with great JS support.

First, install Playwright and browser binaries:

pip install playwright
playwright install chrome

Then the code:

from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    # Launch a headless Chrome browser (set headless=False to see the window)
    browser = p.chromium.launch(headless=True)
    page = browser.new_page()
    page.goto("https://www.sunnxt.com/movie/inside/")
    
    # Extract the token directly from the rendered DOM
    csrf_token = page.locator('meta[name="csrf-token"]').get_attribute("content")
    print("Extracted CSRF Token:", csrf_token)
    
    browser.close()

Solution 3: Scrapy with Playwright Middleware

If you’re using Scrapy for scraping, you can enable JS rendering with the scrapy-playwright middleware:

  1. Install the package:
pip install scrapy-playwright
  1. Update your settings.py to enable the middleware:
DOWNLOAD_HANDLERS = {
    "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
    "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
}

PLAYWRIGHT_LAUNCH_OPTIONS = {
    "headless": True,
}
  1. Modify your spider to use JS rendering:
import scrapy

class SunnxtSpider(scrapy.Spider):
    name = "sunnxt"
    start_urls = ["https://www.sunnxt.com/movie/inside/"]

    def start_requests(self):
        for url in self.start_urls:
            # Tell Scrapy to use Playwright for this request
            yield scrapy.Request(url, meta={"playwright": True})

    def parse(self, response):
        # Your original CSS selector will now work!
        csrf_token = response.css('meta[name="csrf-token"]::attr(content)').get()
        yield {"csrf_token": csrf_token}

Any of these approaches should let you successfully extract the CSRF token. Let me know if you run into any snags! 😊

内容的提问来源于stack exchange,提问作者Pradeep

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 03:50:04