You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Google Apps Script爬取网站数据遇404错误求助

Troubleshooting 404 Error in Google Apps Script Web Scraping with Login

Let's break down why you're hitting that 404 error and how to fix it—your approach is on the right track, but there are a few key gaps with how modern web apps handle routing and authentication that you're missing:

1. You're requesting frontend routes, not backend APIs

The URLs you're using (https://net.statev.de/#/login and https://net.statev.de/#/pages/buisness/storage/...) include a # (hash), which is a client-side frontend route. When you send a request to these URLs via UrlFetchApp, the server only processes the part before the #—so you're actually asking for https://net.statev.de/ every time, not the login or storage data endpoints.

Modern single-page apps (SPAs) use these hash routes to load content dynamically in the browser, but real authentication and data fetching happen via separate backend API endpoints. You need to find these actual API URLs first:

  • Open your browser's DevTools (F12), go to the Network tab
  • Manually log in, then look for XHR/fetch requests (not document requests) — this will show you the real login API URL (e.g., something like https://net.statev.de/api/auth/login)
  • Repeat for the storage page: navigate to it and find the API call that pulls the data you need.

2. Your login request format is likely incorrect

Most modern APIs expect JSON payloads, but your current code sends data as application/x-www-form-urlencoded (the default for UrlFetchApp when passing an object as payload). If the backend expects JSON, this will fail silently, leaving you with invalid cookies for your subsequent request.

Fix this by stringifying your payload and setting the correct Content-Type header:

var loginPayload = JSON.stringify({ 
  'email':'testmail', 
  'password':"test" 
});
var loginOptions = {
  'method':'post',
  'payload': loginPayload,
  'headers': {
    'Content-Type': 'application/json',
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' // Mimic a browser to avoid bot detection
  },
  'followredirects': false
};

3. You're not handling all required cookies

Your code only extracts the first cookie from the Set-Cookie header, but many apps use multiple cookies for authentication (e.g., session ID, CSRF tokens). Missing any of these will cause your data request to be rejected.

Update your cookie handling to capture all cookies:

var setCookies = loginResponse.getAllHeaders()["Set-Cookie"];
var cookie;
if (Array.isArray(setCookies)) {
  // If multiple cookies, extract each one before the first semicolon
  cookie = setCookies.map(c => c.split(";")[0]).join("; ");
} else {
  cookie = setCookies.split(";")[0];
}

4. Don't forget CSRF tokens (if required)

Many sites use CSRF tokens to block unauthorized requests. Before sending the login request, you may need to fetch the login page's HTML and extract a CSRF token from a meta tag (e.g., <meta name="csrf-token" content="abc123">), then include that token in your login headers (like X-CSRF-Token: abc123).

Final Notes

Once you have the correct API URLs, adjusted payload format, and proper cookie handling, your request should stop returning 404s. Keep in mind that some sites may have additional anti-scraping measures (like rate limiting), so mimicking browser headers can help avoid being blocked.

内容的提问来源于stack exchange,提问作者Torba_Trike

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 08:57:37