如何使用requests无需硬编码Cookie爬取网页表格内容?
Hey there! Let's break down how to solve this problem. The issue here is that the website uses session cookies and anti-bot cookies to validate requests—when you send a request without these, even though you get a 200 OK, the server doesn't serve up the actual table content. Instead of hardcoding cookies (which expire quickly), we can use requests.Session() to automatically handle cookie persistence, just like a web browser does.
Here's the Fix
The requests.Session() object keeps track of cookies set by the server across multiple requests. We'll first send an initial request to let the server generate the necessary cookies, then use the same session to fetch the page again with those cookies included.
Modified Code
import requests from bs4 import BeautifulSoup # Initialize a session to manage cookies automatically session = requests.Session() # Use a valid User-Agent to mimic a real browser headers = { "User-Agent": "Mozilla/5.0 (Windows NT 6.1) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/86.0.4240.183 Safari/537.36" } # First request: let the server set session and anti-bot cookies initial_request = session.get( 'https://www.health.gov.il/Subjects/KidsAndMatures/child_development/Pages/ADHD_experts.aspx', headers=headers ) print(f"Initial request status code: {initial_request.status_code}") # Second request: use the same session (now with valid cookies) to get the table target_request = session.get( 'https://www.health.gov.il/Subjects/KidsAndMatures/child_development/Pages/ADHD_experts.aspx', headers=headers ) print(f"Target request status code: {target_request.status_code}") # Parse and extract the table soup = BeautifulSoup(target_request.text, "lxml") table = soup.select_one('table:has(> caption.resultsSummaryPhones)') print(table)
Why This Works
- The session automatically stores cookies like
ASP.NET_SessionIdandBotMitigationCookiethat the server sends back after the initial request. - When we make the second request, the session includes these cookies in the headers, so the server recognizes our request as valid and serves the full table content.
Extra Notes
- Keep your
User-Agentupdated to match a real browser—some sites block requests with outdated or generic User-Agents. - If the site adds more complex anti-bot measures (like JavaScript-generated cookies), you might need tools like Selenium or Playwright to fully mimic a browser. But for this specific page, the session approach should work perfectly.
内容的提问来源于stack exchange,提问作者MITHU

