You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Sheypoor波斯房产网位置标签抓取及443连接报错修复咨询

问题1:443端口连接超时解决方案

该错误本质是你当前运行环境(Google Colab)无法正常连通目标站点,通常由站点反爬拦截、地域访问限制两类原因导致,可按以下方案排查解决:

  • 配置真实浏览器请求头:默认的爬虫请求头会被大部分站点直接拦截,请求时需携带模拟桌面端浏览器的UA标识
  • 配置代理访问:如果站点存在地域访问限制,可在请求参数中加入可用的HTTPS代理
  • 配置重试机制:增加自动重试策略避免偶发网络波动导致的请求失败
问题2:房产位置信息提取方案

从你提供的示例HTML结构来看,房产位置对应div.content容器下无任何子标签(无time、无attr-label类span)的p标签,推荐用BeautifulSoup做DOM解析,比正则匹配更稳定,定位逻辑如下:

  1. 先定位所有class为content的div节点,每个节点对应一条房产信息
  2. 遍历节点下所有p标签,过滤掉带time子标签的发布时间p、带attr-label子标签的价格p,剩余p标签的文本即为房产位置

调整后完整可用代码示例

import pandas as pd
import requests
from bs4 import BeautifulSoup
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry

# 初始化数据表,新增位置列
real_estate_df = pd.DataFrame(columns=["URL", "نوع ملک","اجاره", "موقعیت"])
urls = ['https://www.sheypoor.com/%D8%A7%DB%8C%D8%B1%D8%A7%D9%86/%D8%A7%D9%85%D9%84%D8%A7%DA%A9/%D8%B1%D9%87%D9%86-%D8%A7%D8%AC%D8%A7%D8%B1%D9%87-%D8%AE%D8%A7%D9%86%D9%87-%D8%A2%D9%BE%D8%A7%D8%B1%D8%AA%D9%85%D8%A7%D9%86']

# 配置请求会话,内置重试逻辑
session = requests.Session()
retry_strategy = Retry(
    total=3,
    backoff_factor=1,
    status_forcelist=[429, 500, 502, 503, 504]
)
adapter = HTTPAdapter(max_retries=retry_strategy)
session.mount("https://", adapter)
session.mount("http://", adapter)

# 模拟浏览器请求头
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
    'Accept-Language': 'fa-IR,fa;q=0.9,en-US;q=0.8,en;q=0.7'
}
# 如有代理可取消注释下行配置
# proxies = {'https': '你的代理服务地址'}

def get_data(url):
    res_list = []
    resp = session.get(url, headers=headers, timeout=10)
    # 如有代理请将上行改为 resp = session.get(url, headers=headers, proxies=proxies, timeout=10)
    resp.encoding = 'utf-8'
    soup = BeautifulSoup(resp.text, 'lxml')
    for item in soup.select('div.content'):
        # 提取位置信息
        location = None
        p_tags = item.find_all('p')
        for p in p_tags:
            if not p.find('time') and not p.find('span', class_='attr-label'):
                location = p.get_text(strip=True)
                break
        # 提取租金信息
        rent_tag = item.find('span', string='اجاره:')
        rent = rent_tag.find_next('strong').get_text(strip=True) if rent_tag else None
        # 提取详情页链接
        detail_url = item.select_one('h2 a')['href'] if item.select_one('h2 a') else None
        res_list.append({
            'URL': detail_url,
            'اجاره': rent,
            'موقعیت': location
            # 其他字段可参照上述逻辑自行补充
        })
    return res_list

for url in urls:
    print(url)
    if 'sheypoor' in url:
        items = get_data(url)
        for item in items:
            # 原有append方法已在新版pandas废弃,改用concat做行拼接
            real_estate_df = pd.concat([real_estate_df, pd.DataFrame([item])], ignore_index=True)

print(real_estate_df.head())

注意:爬取网站数据前请确认符合目标站点的robots协议和相关使用条款,控制请求频率避免对站点服务造成影响。

内容的提问来源于stack exchange,提问作者Soheil Paper

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 23:30:00