使用curl多次访问www.seloger.com遭遇HTTP/1.1 307临时重定向问题
curl Requests to Seloger.com Hey, I’ve run into similar anti-scraping issues with real estate sites like Seloger before—they’re pretty strict about identifying non-human or inconsistent requests. Let’s break down how to fix this:
1. Preserve Session Cookies Between Requests
The first request probably sets session cookies that the server expects you to send on subsequent visits. If you don’t include them, the server flags your request as suspicious and redirects you. Here’s how to handle it:
- Save cookies from your first successful request:
curl -i -c seloger_cookies.txt https://www.seloger.com/your-target-url - Reuse those cookies for all follow-up requests:
curl -i -b seloger_cookies.txt https://www.seloger.com/your-target-url
2. Mimic a Real Browser’s Request Headers
Curl’s default headers are super generic, and Seloger’s servers can spot them easily. Add headers that match what a modern browser sends:
curl -i \ -b seloger_cookies.txt \ -H "User-Agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" \ -H "Accept: text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8" \ -H "Accept-Language: en-US,en;q=0.5" \ -H "Connection: keep-alive" \ -H "Upgrade-Insecure-Requests: 1" \ https://www.seloger.com/your-target-url
This combination of cookies and browser-like headers should make your requests look far more legitimate.
3. Check for Rate Limiting
Even if you switch WiFi, Seloger might be tracking request patterns (like frequent hits from the same User-Agent, or rapid consecutive requests). Try adding a delay between requests:
sleep 5 && curl -i -b seloger_cookies.txt -H "User-Agent: ..." https://www.seloger.com/your-target-url
You can also tweak the User-Agent slightly between requests (e.g., change the Chrome version to 117.0.0.0) to avoid being flagged as a bot.
4. Inspect the Redirected Error Page
Use the -L flag to follow the 307 redirect and see what the error page says—this can give you clues about why you’re being blocked:
curl -L -i https://www.seloger.com/your-target-url
If the page mentions a CAPTCHA or verification, you might need a more advanced tool like Playwright/Puppeteer to simulate a full browser environment, but for basic curl use cases, the first three steps should resolve the issue.
内容的提问来源于stack exchange,提问作者rebzy

