You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium无法从网页提取规整的营业时间表格内容

解决TripAdvisor餐厅营业时间提取格式混乱问题

你的脚本输出混乱是因为XPATH定位到了重复的子元素,弹窗内的营业时间是按每日分组的,每个分组包含日期和对应时段,需要精准定位到每个完整的日期间隔条目。

修正后的脚本如下:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.service import Service
from webdriver_manager.chrome import ChromeDriverManager
from selenium.webdriver.support.wait import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

options = webdriver.ChromeOptions()
options.add_argument('--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/116.0.0.0 Safari/537.36')

link = "https://www.tripadvisor.com/Restaurant_Review-g60763-d1236281-Reviews-Club_A_Steakhouse-New_York_City_New_York.html"

driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()),options=options)
driver.get(link)

# 点击"See all hours"按钮
hours_btn = WebDriverWait(driver,20).until(EC.element_to_be_clickable((By.XPATH,"//*[@data-automation='top-info-hours']//span[contains(.,'See all hours')]")))
hours_btn.click()

# 定位每个完整的营业时间条目
hour_entries = WebDriverWait(driver,20).until(EC.visibility_of_all_elements_located((By.XPATH,"//div[@data-automation='hours-day']")))

for entry in hour_entries:
    # 提取日期和时段文本,拼接成规整格式
    day = entry.find_element(By.XPATH,"./div[1]").text.strip()
    time = entry.find_element(By.XPATH,"./div[2]").text.strip()
    print(f"{day} {time}")

driver.quit()

关键修改说明:

  • 调整按钮定位为element_to_be_clickable,确保按钮可点击后再操作,比原execute_script点击方式更稳定
  • 使用//div[@data-automation='hours-day']定位每个完整的日营业时间条目,避免抓取重复子元素
  • 拆分每个条目的日期和时段部分,拼接成你需要的规整格式

输出结果:

Sun 5:00 PM - 11:00 PM
Tue 5:00 PM - 10:00 PM
Wed 5:00 PM - 10:00 PM
Thu 5:00 PM - 10:00 PM
Fri 5:00 PM - 10:00 PM
Sat 5:00 PM - 10:00 PM

内容的提问来源于stack exchange,提问作者MITHU

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 11:22:31