如何高效爬取Bayut平台DLD认证房源数据且避免401错误?
爬取Bayut平台房产数据时无法获取DLD认证绿标信息的问题
我正在使用Scrapy爬取Bayut平台的房产数据,但无法提取绿标(DLD认证信息),具体情况如下:
- 该信息通过带基础认证的POST API获取。
- 完全复制network面板中的请求头、负载和参数,在Postman/Python中调用API时返回401 Unauthorized错误。
- 使用Selenium可以成功获取,但每周需爬取约21万条房源,Selenium速度太慢无法满足需求。
已尝试方案:
- Scrapy:无法获取认证信息。
- Postman及Python requests:返回401错误。
- Selenium:可行但速度过慢。
网站是否可能采用了会话认证或IP限制等额外安全措施?如何高效获取这些数据?
附上请求POST API的代码:
import requests import base64 # Define the URL url = "https://fenix-data-es2.bayut.com/_msearch" # Encode credentials manually (decoded: "bayut_read_user_es2:10yNmg5+6K") auth_string = "bayut_read_user_es2:10yNmg5+6K" auth_encoded = base64.b64encode(auth_string.encode()).decode() # Convert to Base64 # Headers with Authorization headers = { "Authorization": f"Basic {auth_encoded}", "accept": "*/*", "accept-encoding": "gzip, deflate, br, zstd", "accept-language": "en-US,en;q=0.9", "cache-control": "no-cache", "content-type": "application/x-ndjson", "origin": "https://www.bayut.com", "pragma": "no-cache", "priority": "u=1, i", "referer": "https://www.bayut.com/", "sec-ch-ua": "\"Not(A:Brand\";v=\"99\", \"Google Chrome\";v=\"133\", \"Chromium\";v=\"133\"", "sec-ch-ua-mobile": "?0", "sec-ch-ua-platform": "\"Windows\"", "sec-fetch-dest": "empty", "sec-fetch-mode": "cors", "sec-fetch-site": "same-site", "user-agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/133.0.0.0 Safari/537.36" } # Query parameters (filter_path) params = { "filter_path": "took,*.took,*.suggest.*.options.text,*.suggest.*.options._source.*,*.hits.total.*,*.hits.hits._source.*,*.hits.hits._score,*.hits.hits.highlight.*,*.error,*.aggregations.*.buckets.key,*.aggregations.*.buckets.doc_count,*.aggregations.*.buckets.complex_value.hits.hits._source,*.aggregations.*.filtered_agg.facet.buckets.key,*.aggregations.*.filtered_agg.facet.buckets.doc_count,*.aggregations.*.filtered_agg.facet.buckets.complex_value.hits.hits._source" } # POST data (formatted in NDJSON format) post_data = """{"index":"dld_matched_property_details_prod_alias"} {"from":0,"size":5,"track_total_hits":10000,"query":{"bool":{"must":[{"term":{"external_id":"10228377"}}]}}} """ # Sending the POST request response = requests.post(url, headers=headers, params=params, data=post_data) # Check if the request was successful if response.status_code == 200: print("Request Successful!") print(response.json()) # Print the response in JSON format else: print(f"Request failed with status code: {response.status_code}") print(response.text) # Print the error message if any
解决方案建议
一、排查401错误的核心原因
- 会话绑定验证:Bayut大概率会在用户访问网站时生成会话Cookie(如
__cf_bm、sessionId这类),API请求必须携带这些会话Cookie才能通过验证。可以在浏览器访问房源页面后,复制Cookie到请求头中测试。 - 签名验证:部分网站会对请求参数或body生成签名(比如基于时间戳+密钥的HMAC签名),需要检查network面板中是否有额外的签名头(如
x-signature、x-timestamp),或请求body中是否有隐藏的签名字段。 - IP/UA动态校验:网站可能会验证请求IP和UA的匹配度,或对陌生IP进行拦截。可以尝试使用代理IP池,同时保持UA和浏览器完全一致,避免频繁切换UA。
- 基础认证时效性:你使用的Basic Auth凭证可能是临时的,会随会话过期失效。可以尝试从浏览器的实时请求头中提取最新的Authorization字段,而非硬编码。
二、高效爬取的优化方案
- Scrapy集成Cookie池+代理池
- 在Scrapy中使用
CookieMiddleware维护会话Cookie,每次请求前从已成功的请求或浏览器中获取有效Cookie。 - 配置代理池(如
scrapy-proxies),分散请求IP,避免被限制。
- 在Scrapy中使用
- 补全请求细节
- 确保请求头中的
sec-ch-ua、sec-ch-ua-platform等字段格式完全和浏览器一致(你的代码中sec-ch-ua曾有多余换行,可能导致格式错误)。 - 添加
Cookie头,包含浏览器中获取的所有相关Cookie(如bayut_session、cf_clearance等)。
- 确保请求头中的
- 批量请求优化
- 该API是
_msearch接口,支持批量查询多个external_id,可以把多个房源ID放在一个请求中,减少请求次数提升效率。示例修改post_data:{"index":"dld_matched_property_details_prod_alias"} {"from":0,"size":5,"track_total_hits":10000,"query":{"bool":{"must":[{"terms":{"external_id":["10228377","10228378","10228379"]}}]}}}
- 该API是
- 用Scrapy+Playwright替代Selenium
- Playwright比Selenium更轻量快速,可以在Scrapy中集成
scrapy-playwright,仅在需要获取会话Cookie或验证信息时启动浏览器,其余请求用常规Scrapy请求,兼顾速度和成功率。
- Playwright比Selenium更轻量快速,可以在Scrapy中集成
三、其他注意事项
- 避免短时间内发送大量请求,设置合理的请求间隔(如
DOWNLOAD_DELAY = 1),或通过Scrapy的CONCURRENT_REQUESTS控制并发数。 - 定期检查API变化,Bayut可能更新认证机制,需及时调整请求参数和头信息。
内容的提问来源于stack exchange,提问作者Shashank Nakka
相关产品推荐
相关产品推荐

