如何用BeautifulSoup从lefooding网站提取餐厅WEBSITE字段链接?
解决方案:提取lefooding餐厅页面的WEBSITE链接
问题分析
你之前的代码没找到目标链接,大概率是两个原因:
- 页面内容是JavaScript动态渲染的,
requests.get()只能获取初始HTML,拿不到JS加载后的内容; - 选择器
e-rowContent不符合页面实际结构,定位不到目标区域。
下面提供两种可行方案,按优先级尝试:
方案1:静态HTML解析(优先尝试)
如果目标内容是静态嵌入的,可通过定位"WEBSITE"文本元素,再关联到对应的链接标签。同时添加User-Agent请求头,避免被网站反爬拦截:
from bs4 import BeautifulSoup import requests url = "https://lefooding.com/en/restaurants/ezkia" # 模拟浏览器请求头 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } res = requests.get(url, headers=headers) soup = BeautifulSoup(res.text, 'html.parser') # 定位包含"WEBSITE"的文本节点 website_label = soup.find(string=lambda text: text and "WEBSITE" in text.strip()) if website_label: # 查找该标签相邻的链接元素 website_link = website_label.find_next("a") if website_link: print("提取到的WEBSITE链接:", website_link.get("href")) else: print("未找到WEBSITE对应的链接") else: print("页面中未找到WEBSITE标签")
方案2:动态渲染页面解析(静态方案失效时用)
如果目标内容是JS动态加载的,需要用selenium模拟浏览器加载完整页面:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.chrome.service import Service from webdriver_manager.chrome import ChromeDriverManager import time url = "https://lefooding.com/en/restaurants/ezkia" # 自动安装并启动Chrome浏览器 driver = webdriver.Chrome(service=Service(ChromeDriverManager().install())) driver.get(url) time.sleep(3) # 等待页面完全加载(可根据网络情况调整) try: # 通过XPath定位WEBSITE文本后的链接 website_link = driver.find_element(By.XPATH, "//*[contains(text(), 'WEBSITE')]/following-sibling::a") print("提取到的WEBSITE链接:", website_link.get_attribute("href")) except Exception as e: print("提取失败:", str(e)) finally: driver.quit()
关键说明
- 静态方案更高效,适合内容未动态渲染的场景;
- 动态方案需要安装
selenium和对应浏览器驱动,执行前先运行pip install selenium webdriver-manager; - 若网站有反爬机制,可能需要添加更多请求头或使用代理。
内容的提问来源于stack exchange,提问作者LJG
相关产品推荐
相关产品推荐

