You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用requests无需硬编码Cookie爬取网页表格内容?

How to Scrape the ADHD Experts Table Without Hardcoding Cookies

Hey there! Let's break down how to solve this problem. The issue here is that the website uses session cookies and anti-bot cookies to validate requests—when you send a request without these, even though you get a 200 OK, the server doesn't serve up the actual table content. Instead of hardcoding cookies (which expire quickly), we can use requests.Session() to automatically handle cookie persistence, just like a web browser does.

Here's the Fix

The requests.Session() object keeps track of cookies set by the server across multiple requests. We'll first send an initial request to let the server generate the necessary cookies, then use the same session to fetch the page again with those cookies included.

Modified Code

import requests
from bs4 import BeautifulSoup

# Initialize a session to manage cookies automatically
session = requests.Session()

# Use a valid User-Agent to mimic a real browser
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 6.1) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/86.0.4240.183 Safari/537.36"
}

# First request: let the server set session and anti-bot cookies
initial_request = session.get(
    'https://www.health.gov.il/Subjects/KidsAndMatures/child_development/Pages/ADHD_experts.aspx',
    headers=headers
)
print(f"Initial request status code: {initial_request.status_code}")

# Second request: use the same session (now with valid cookies) to get the table
target_request = session.get(
    'https://www.health.gov.il/Subjects/KidsAndMatures/child_development/Pages/ADHD_experts.aspx',
    headers=headers
)
print(f"Target request status code: {target_request.status_code}")

# Parse and extract the table
soup = BeautifulSoup(target_request.text, "lxml")
table = soup.select_one('table:has(> caption.resultsSummaryPhones)')
print(table)

Why This Works

  • The session automatically stores cookies like ASP.NET_SessionId and BotMitigationCookie that the server sends back after the initial request.
  • When we make the second request, the session includes these cookies in the headers, so the server recognizes our request as valid and serves the full table content.

Extra Notes

  • Keep your User-Agent updated to match a real browser—some sites block requests with outdated or generic User-Agents.
  • If the site adds more complex anti-bot measures (like JavaScript-generated cookies), you might need tools like Selenium or Playwright to fully mimic a browser. But for this specific page, the session approach should work perfectly.

内容的提问来源于stack exchange,提问作者MITHU

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 20:37:37