You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python的requests和BeautifulSoup4解析指定网站的非精选广告?

解决方案:提取普通租房广告(排除精选/存档)

针对你遇到的广告类名不统一、精选/存档广告与普通广告共用父类、多页结构不一致的问题,可通过识别广告的独特标识精准过滤目标内容,具体实现如下:

核心思路

不依赖统一类名或“Lista ogłoszeń”字段,直接遍历所有广告容器,通过以下特征排除非目标广告:

  • 精选广告:通常带有专属标识(比如页面内的wybrane-ogloszenie类、“Wybrane”标签)
  • 存档广告:会标注“Archiwum”文本,或带有存档相关属性/类

代码实现

import requests
from bs4 import BeautifulSoup

# 示例页面URL,分页时替换page参数即可
url = "https://www.nieruchomosci-online.pl/szukaj.html?3,mieszkanie,wynajem,,Szczecin:19503"
# 模拟浏览器请求头,避免反爬拦截
headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"}

response = requests.get(url, headers=headers)
soup = BeautifulSoup(response.text, 'html.parser')

# 获取所有广告容器
all_ad_containers = soup.find_all(class_='column-container column_default')

normal_ads = []
for container in all_ad_containers:
    # 检测精选广告:根据页面实际标识调整,比如查找带"wybrane-ogloszenie"的元素
    is_promoted = container.find(class_='wybrane-ogloszenie') is not None
    # 检测存档广告:查找包含"Archiwum"的文本
    is_archived = container.find(string=lambda text: text and 'Archiwum' in text.strip()) is not None
    
    # 仅保留普通广告
    if not is_promoted and not is_archived:
        normal_ads.append(container)

# 提取普通广告关键信息(示例)
for idx, ad in enumerate(normal_ads, 1):
    title = ad.find('h2').get_text(strip=True) if ad.find('h2') else '无标题'
    price = ad.find(class_='price').get_text(strip=True) if ad.find(class_='price') else '无价格'
    print(f"普通广告 {idx}:")
    print(f"  标题: {title}")
    print(f"  价格: {price}\n")

注意事项

  1. 若页面结构更新,需重新检查精选/存档广告的独特特征:
    • 右键查看精选广告的HTML,确认是否有专属类名(如promoted、premium)
    • 存档广告可能带有data-archived="true"这类属性,或文本包含“Archiwum”
  2. 分页处理时,只需在URL后添加&page=2这类参数,重复上述逻辑即可

内容的提问来源于stack exchange,提问作者Andrey

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 04:07:21