使用Google Apps Script爬取网站数据遇404错误求助
Let's break down why you're hitting that 404 error and how to fix it—your approach is on the right track, but there are a few key gaps with how modern web apps handle routing and authentication that you're missing:
1. You're requesting frontend routes, not backend APIs
The URLs you're using (https://net.statev.de/#/login and https://net.statev.de/#/pages/buisness/storage/...) include a # (hash), which is a client-side frontend route. When you send a request to these URLs via UrlFetchApp, the server only processes the part before the #—so you're actually asking for https://net.statev.de/ every time, not the login or storage data endpoints.
Modern single-page apps (SPAs) use these hash routes to load content dynamically in the browser, but real authentication and data fetching happen via separate backend API endpoints. You need to find these actual API URLs first:
- Open your browser's DevTools (F12), go to the Network tab
- Manually log in, then look for XHR/fetch requests (not document requests) — this will show you the real login API URL (e.g., something like
https://net.statev.de/api/auth/login) - Repeat for the storage page: navigate to it and find the API call that pulls the data you need.
2. Your login request format is likely incorrect
Most modern APIs expect JSON payloads, but your current code sends data as application/x-www-form-urlencoded (the default for UrlFetchApp when passing an object as payload). If the backend expects JSON, this will fail silently, leaving you with invalid cookies for your subsequent request.
Fix this by stringifying your payload and setting the correct Content-Type header:
var loginPayload = JSON.stringify({ 'email':'testmail', 'password':"test" }); var loginOptions = { 'method':'post', 'payload': loginPayload, 'headers': { 'Content-Type': 'application/json', 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' // Mimic a browser to avoid bot detection }, 'followredirects': false };
3. You're not handling all required cookies
Your code only extracts the first cookie from the Set-Cookie header, but many apps use multiple cookies for authentication (e.g., session ID, CSRF tokens). Missing any of these will cause your data request to be rejected.
Update your cookie handling to capture all cookies:
var setCookies = loginResponse.getAllHeaders()["Set-Cookie"]; var cookie; if (Array.isArray(setCookies)) { // If multiple cookies, extract each one before the first semicolon cookie = setCookies.map(c => c.split(";")[0]).join("; "); } else { cookie = setCookies.split(";")[0]; }
4. Don't forget CSRF tokens (if required)
Many sites use CSRF tokens to block unauthorized requests. Before sending the login request, you may need to fetch the login page's HTML and extract a CSRF token from a meta tag (e.g., <meta name="csrf-token" content="abc123">), then include that token in your login headers (like X-CSRF-Token: abc123).
Final Notes
Once you have the correct API URLs, adjusted payload format, and proper cookie handling, your request should stop returning 404s. Keep in mind that some sites may have additional anti-scraping measures (like rate limiting), so mimicking browser headers can help avoid being blocked.
内容的提问来源于stack exchange,提问作者Torba_Trike

