如何用Python的requests和BeautifulSoup4解析指定网站的非精选广告?
解决方案:提取普通租房广告(排除精选/存档)
针对你遇到的广告类名不统一、精选/存档广告与普通广告共用父类、多页结构不一致的问题,可通过识别广告的独特标识精准过滤目标内容,具体实现如下:
核心思路
不依赖统一类名或“Lista ogłoszeń”字段,直接遍历所有广告容器,通过以下特征排除非目标广告:
- 精选广告:通常带有专属标识(比如页面内的
wybrane-ogloszenie类、“Wybrane”标签) - 存档广告:会标注“Archiwum”文本,或带有存档相关属性/类
代码实现
import requests from bs4 import BeautifulSoup # 示例页面URL,分页时替换page参数即可 url = "https://www.nieruchomosci-online.pl/szukaj.html?3,mieszkanie,wynajem,,Szczecin:19503" # 模拟浏览器请求头,避免反爬拦截 headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"} response = requests.get(url, headers=headers) soup = BeautifulSoup(response.text, 'html.parser') # 获取所有广告容器 all_ad_containers = soup.find_all(class_='column-container column_default') normal_ads = [] for container in all_ad_containers: # 检测精选广告:根据页面实际标识调整,比如查找带"wybrane-ogloszenie"的元素 is_promoted = container.find(class_='wybrane-ogloszenie') is not None # 检测存档广告:查找包含"Archiwum"的文本 is_archived = container.find(string=lambda text: text and 'Archiwum' in text.strip()) is not None # 仅保留普通广告 if not is_promoted and not is_archived: normal_ads.append(container) # 提取普通广告关键信息(示例) for idx, ad in enumerate(normal_ads, 1): title = ad.find('h2').get_text(strip=True) if ad.find('h2') else '无标题' price = ad.find(class_='price').get_text(strip=True) if ad.find(class_='price') else '无价格' print(f"普通广告 {idx}:") print(f" 标题: {title}") print(f" 价格: {price}\n")
注意事项
- 若页面结构更新,需重新检查精选/存档广告的独特特征:
- 右键查看精选广告的HTML,确认是否有专属类名(如
promoted、premium) - 存档广告可能带有
data-archived="true"这类属性,或文本包含“Archiwum”
- 右键查看精选广告的HTML,确认是否有专属类名(如
- 分页处理时,只需在URL后添加
&page=2这类参数,重复上述逻辑即可
内容的提问来源于stack exchange,提问作者Andrey
相关产品推荐
相关产品推荐

