使用User-Agent爬取Justdial仍遇HTTPError: Forbidden问题求助
解决Justdial爬取时的403 Forbidden错误
嘿,我懂你现在碰到的麻烦——只设置User-Agent确实对付不了Justdial这类反爬严格的网站,它们会检查一堆请求细节来识别爬虫。给你几个实用的调整方案,一步步搞定这个403问题:
1. 补全请求头,模拟真实浏览器行为
光有User-Agent远远不够,真实浏览器发送请求时会带一堆额外头信息,我们得把这些补上,让请求看起来更“像人”:
import urllib.request from bs4 import BeautifulSoup import urllib.error # 贴近真实Chrome浏览器的完整请求头 headers = { 'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8', 'Accept-Language': 'en-US,en;q=0.5', 'Referer': 'https://www.justdial.com/', # 模拟从首页跳转至目标页 'DNT': '1', # 标记为不允许追踪 'Connection': 'keep-alive', 'Upgrade-Insecure-Requests': '1' } target_url = "你要爬取的Justdial具体页面URL" req = urllib.request.Request(target_url, headers=headers) try: response = urllib.request.urlopen(req) # 解析HTML内容 soup = BeautifulSoup(response.read(), 'html.parser') # 这里写你的CSV解析与生成逻辑 except urllib.error.HTTPError as e: print(f"请求失败: 错误代码 {e.code}, 原因: {e.reason}")
2. 维持会话状态(自动处理Cookie)
很多网站会用Cookie识别合法会话,我们可以用CookieJar自动管理Cookie,模拟浏览器的会话持久化:
import urllib.request from http.cookiejar import CookieJar from bs4 import BeautifulSoup # 创建CookieJar并绑定到请求器 cj = CookieJar() opener = urllib.request.build_opener(urllib.request.HTTPCookieProcessor(cj)) urllib.request.install_opener(opener) # 先访问首页获取初始会话Cookie opener.open("https://www.justdial.com/") # 再发送目标请求,自动带上Cookie和完整请求头 headers = { # 同上的完整请求头 } req = urllib.request.Request(target_url, headers=headers) response = opener.open(req) soup = BeautifulSoup(response.read(), 'html.parser')
3. 添加请求延迟,降低访问频率
Justdial会检测请求频率,短时间内密集请求很容易被封禁。每次请求后添加几秒延迟:
import time # 每次请求完成后延迟3秒,可根据情况调整 time.sleep(3)
4. 备选方案:用Selenium模拟真实浏览器
如果上面的方法还是不行,说明网站可能用了JS动态渲染内容或者更严格的反爬机制。这时候可以用Selenium直接操控浏览器,完全模拟人类操作:
from selenium import webdriver from selenium.webdriver.chrome.options import Options import time from bs4 import BeautifulSoup # 配置Chrome浏览器参数 chrome_options = Options() chrome_options.add_argument("--headless=new") # 无头模式,不显示浏览器窗口 chrome_options.add_argument("user-agent=Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36") driver = webdriver.Chrome(options=chrome_options) driver.get(target_url) time.sleep(3) # 等待页面完全加载 # 获取渲染后的页面HTML page_source = driver.page_source soup = BeautifulSoup(page_source, 'html.parser') # 后续的解析与CSV生成逻辑 driver.quit()
最后提一句:爬取前一定要查看Justdial的robots.txt和服务条款,确保你的爬取行为符合规定,避免不必要的法律风险哦。
内容的提问来源于stack exchange,提问作者prashant dhyani
相关产品推荐
相关产品推荐

