You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用User-Agent爬取Justdial仍遇HTTPError: Forbidden问题求助

解决Justdial爬取时的403 Forbidden错误

嘿,我懂你现在碰到的麻烦——只设置User-Agent确实对付不了Justdial这类反爬严格的网站,它们会检查一堆请求细节来识别爬虫。给你几个实用的调整方案,一步步搞定这个403问题:

1. 补全请求头,模拟真实浏览器行为

光有User-Agent远远不够,真实浏览器发送请求时会带一堆额外头信息,我们得把这些补上,让请求看起来更“像人”:

import urllib.request
from bs4 import BeautifulSoup
import urllib.error

# 贴近真实Chrome浏览器的完整请求头
headers = {
    'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36',
    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8',
    'Accept-Language': 'en-US,en;q=0.5',
    'Referer': 'https://www.justdial.com/',  # 模拟从首页跳转至目标页
    'DNT': '1',  # 标记为不允许追踪
    'Connection': 'keep-alive',
    'Upgrade-Insecure-Requests': '1'
}

target_url = "你要爬取的Justdial具体页面URL"
req = urllib.request.Request(target_url, headers=headers)

try:
    response = urllib.request.urlopen(req)
    # 解析HTML内容
    soup = BeautifulSoup(response.read(), 'html.parser')
    # 这里写你的CSV解析与生成逻辑
except urllib.error.HTTPError as e:
    print(f"请求失败: 错误代码 {e.code}, 原因: {e.reason}")

2. 维持会话状态(自动处理Cookie)

很多网站会用Cookie识别合法会话,我们可以用CookieJar自动管理Cookie,模拟浏览器的会话持久化:

import urllib.request
from http.cookiejar import CookieJar
from bs4 import BeautifulSoup

# 创建CookieJar并绑定到请求器
cj = CookieJar()
opener = urllib.request.build_opener(urllib.request.HTTPCookieProcessor(cj))
urllib.request.install_opener(opener)

# 先访问首页获取初始会话Cookie
opener.open("https://www.justdial.com/")

# 再发送目标请求,自动带上Cookie和完整请求头
headers = {
    # 同上的完整请求头
}
req = urllib.request.Request(target_url, headers=headers)
response = opener.open(req)
soup = BeautifulSoup(response.read(), 'html.parser')

3. 添加请求延迟,降低访问频率

Justdial会检测请求频率,短时间内密集请求很容易被封禁。每次请求后添加几秒延迟:

import time

# 每次请求完成后延迟3秒,可根据情况调整
time.sleep(3)

4. 备选方案:用Selenium模拟真实浏览器

如果上面的方法还是不行,说明网站可能用了JS动态渲染内容或者更严格的反爬机制。这时候可以用Selenium直接操控浏览器,完全模拟人类操作:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
import time
from bs4 import BeautifulSoup

# 配置Chrome浏览器参数
chrome_options = Options()
chrome_options.add_argument("--headless=new")  # 无头模式,不显示浏览器窗口
chrome_options.add_argument("user-agent=Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36")

driver = webdriver.Chrome(options=chrome_options)
driver.get(target_url)
time.sleep(3)  # 等待页面完全加载

# 获取渲染后的页面HTML
page_source = driver.page_source
soup = BeautifulSoup(page_source, 'html.parser')

# 后续的解析与CSV生成逻辑
driver.quit()

最后提一句:爬取前一定要查看Justdial的robots.txt和服务条款,确保你的爬取行为符合规定,避免不必要的法律风险哦。

内容的提问来源于stack exchange,提问作者prashant dhyani

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:25:06