使用Simple Html Dom爬取Booking.com:仅输出含/hotel/的结果链接及标题
解决方法:提取Booking.com的酒店链接和标题
嘿,我来帮你搞定这个问题!你现在需要筛选出包含/hotel/路径的链接,同时提取对应的酒店标题,下面是具体的调整方案和代码说明:
核心思路
- 精准筛选链接:用正则匹配包含
/hotel/路径的完整HTTP/HTTPS链接 - 提取对应标题:Booking的酒店结果都有专门的容器和标题元素,直接定位这些元素比遍历所有
<a>标签更可靠 - 处理相对路径:部分链接可能是相对路径,需要补全域名确保链接可用
修改后的完整代码
<?php include 'simple_html_dom.php'; $searchText = "Venice"; $searchText = str_replace(" ", "+", $searchText); $url = "https://www.booking.com/searchresults.en-gb.html?aid=1781605&lang=en-gb&sid=3bb432f656e368125330f71ea0e74e36&sb=1&src=index&src_elem=sb&error_url=https://www.booking.com/index.en-gb.html?aid=1781605;sid=3bb432f656e368125330f71ea0e74e36;sb_price_type=total;srpvid=dc2798d544dd007f&;&ss=".$searchText."&is_ski_area=0&ssne=".$searchText."&ssne_untouched=".$searchText."&dest_id=-132007&dest_type=city&checkin_year=2019&checkin_month=5&checkin_monthday=19&checkout_year=2019&checkout_month=5&checkout_monthday=20&group_adults=2&group_children=0&no_rooms=1&b_h4u_keep_filters=&from_sf=1"; print $url."<br>"; $html = file_get_html($url); $i = 0; // 直接定位酒店结果容器,避免匹配无关链接 $hotelContainers = $html->find('div.sr_item'); foreach ($hotelContainers as $container) { // 提取酒店名称链接(Booking页面专门用于酒店标题的元素) $linkObj = $container->find('a.hotel_name_link', 0); if (!$linkObj) continue; $link = trim($linkObj->href); // 补全相对路径为完整链接 if (!preg_match('/^https?:\/\//', $link)) { $link = 'https://www.booking.com' . $link; } // 提取酒店标题(纯文本格式) $title = trim($linkObj->plaintext); $i++; echo "[$i] 酒店标题: " . $title . "<br>"; echo "酒店链接: " . $link . "<br/><hr/>"; } ?>
关键细节说明
- 精准定位容器:通过
div.sr_item定位每个酒店的结果容器,这样不会匹配到页脚、导航栏里的无关/hotel/链接 - 链接处理:如果遇到相对路径(比如
/hotel/it/xxx.html),自动补全Booking的域名,确保链接可以直接访问 - 标题提取:使用
a.hotel_name_link元素获取标题,这是Booking页面中专门用于展示酒店名称的标签,比遍历所有<a>标签更可靠
重要提醒
Booking.com有反爬机制,频繁请求可能会导致IP被封禁,建议添加请求间隔(比如用sleep()函数),同时遵守网站的robots.txt规则。如果后续页面结构变化,需要重新检查元素的类名或结构哦。
内容的提问来源于stack exchange,提问作者Rudra Sarkar
相关产品推荐
相关产品推荐

