使用Python+BeautifulSoup爬取TripAdvisor时因分页无法获取全部餐厅链接
解决TripAdvisor马德里餐厅分页爬取问题
你的问题出在仅请求了第一页数据,未处理TripAdvisor的分页逻辑。TripAdvisor通过-oaXX参数控制分页偏移量,XX代表跳过的条目数(每页固定30条,因此第二页偏移量为30,第三页为60,以此类推)。
以下是修改后的代码,支持自动遍历所有分页并提取全部餐厅链接:
import requests from bs4 import BeautifulSoup import time city_name = 'Madrid' geo_code = '187514' base_url = "https://www.tripadvisor.com/Restaurants-g{}-{}-oa{}.html" headers = { "user-agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36" } # 打开文件保存链接,自动管理资源 with open('madrid_restaurants.txt', 'w', encoding='utf-8') as f: offset = 0 while True: # 构造当前分页的请求URL url = base_url.format(geo_code, city_name, offset) response = requests.get(url, headers=headers) soup = BeautifulSoup(response.text, 'html.parser') # 用CSS选择器提取餐厅链接,比正则更稳定 restaurant_links = soup.select('a[data-test="restaurant-name"]') if not restaurant_links: # 无更多链接,终止循环 break # 遍历提取到的链接,补全完整URL并保存 for link in restaurant_links: full_link = "https://www.tripadvisor.com" + link['href'] print(full_link) f.write(f"{full_link}\n") # 检查是否存在下一页按钮 next_page_btn = soup.select_one('a.nav.next') if not next_page_btn: break # 更新偏移量,准备请求下一页 offset += 30 # 添加延时,避免请求过于频繁触发反爬 time.sleep(2)
核心优化说明:
- 分页循环控制:通过
offset变量递增30,实现分页URL的动态生成 - 稳定链接提取:使用BeautifulSoup的CSS选择器替代正则表达式,适配页面结构变化的能力更强
- 循环终止条件:当无法提取到餐厅链接或无下一页按钮时,自动停止爬取
- 反爬防护:加入
time.sleep(2)降低请求频率,减少被网站封禁的风险 - 文件操作优化:采用
with语句自动管理文件资源,避免遗漏关闭操作
内容的提问来源于stack exchange,提问作者MDON
相关产品推荐
相关产品推荐

