无法通过CSS/XPath提取网页CSRF Token,请求技术协助
Hey there! The selectors you’re using (meta[name="csrf-token"]::attr(content) for CSS and //meta[@name="csrf-token"]/@content for XPath) are actually perfectly correct—so the issue is almost certainly that the page loads this meta tag dynamically using JavaScript. Static HTML-fetching tools (like Python’s requests library) only grab the initial server response, which doesn’t include elements injected after the page loads via client-side code.
Here’s a breakdown of the problem and fixes:
Why Your Current Approach Isn’t Working
If you right-click the page and select "View Page Source", you probably won’t see the csrf-token meta tag listed there. But if you use your browser’s dev tools (F12) to inspect the DOM, you will see it. That’s a clear sign the token is added dynamically by JavaScript after the initial HTML loads. Tools that don’t execute JS can’t access elements added this way.
Solutions Using JavaScript-Rendering Tools
You’ll need to use a tool that simulates a real browser (and runs JS) to access the fully rendered DOM. Here are a few popular options with code examples:
Solution 1: Python + Selenium
Selenium controls a real browser, so it renders the page exactly like a human would.
First, make sure you have Selenium and ChromeDriver installed (or use another browser like Firefox):
pip install selenium # Download ChromeDriver from https://sites.google.com/chromium.org/driver/ and add it to your PATH
Then the code:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # Initialize the browser driver = webdriver.Chrome() try: driver.get("https://www.sunnxt.com/movie/inside/") # Wait up to 10 seconds for the meta tag to appear in the DOM csrf_meta = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CSS_SELECTOR, 'meta[name="csrf-token"]')) ) csrf_token = csrf_meta.get_attribute("content") print("Extracted CSRF Token:", csrf_token) finally: # Always close the browser when done driver.quit()
Solution 2: Python + Playwright
Playwright is a modern, fast alternative to Selenium with great JS support.
First, install Playwright and browser binaries:
pip install playwright playwright install chrome
Then the code:
from playwright.sync_api import sync_playwright with sync_playwright() as p: # Launch a headless Chrome browser (set headless=False to see the window) browser = p.chromium.launch(headless=True) page = browser.new_page() page.goto("https://www.sunnxt.com/movie/inside/") # Extract the token directly from the rendered DOM csrf_token = page.locator('meta[name="csrf-token"]').get_attribute("content") print("Extracted CSRF Token:", csrf_token) browser.close()
Solution 3: Scrapy with Playwright Middleware
If you’re using Scrapy for scraping, you can enable JS rendering with the scrapy-playwright middleware:
- Install the package:
pip install scrapy-playwright
- Update your
settings.pyto enable the middleware:
DOWNLOAD_HANDLERS = { "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", } PLAYWRIGHT_LAUNCH_OPTIONS = { "headless": True, }
- Modify your spider to use JS rendering:
import scrapy class SunnxtSpider(scrapy.Spider): name = "sunnxt" start_urls = ["https://www.sunnxt.com/movie/inside/"] def start_requests(self): for url in self.start_urls: # Tell Scrapy to use Playwright for this request yield scrapy.Request(url, meta={"playwright": True}) def parse(self, response): # Your original CSS selector will now work! csrf_token = response.css('meta[name="csrf-token"]::attr(content)').get() yield {"csrf_token": csrf_token}
Any of these approaches should let you successfully extract the CSRF token. Let me know if you run into any snags! 😊
内容的提问来源于stack exchange,提问作者Pradeep

