You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从Facebook页面下拉菜单提取格式化的营业时间数据

解决Facebook页面营业时间爬取格式混乱问题

你的问题出在直接提取了整个aria-label="Hours"元素的文本,导致包含了热门时段、当前状态等无关内容。要获取规整的「星期+营业时间」格式,需要定位到每个单独的星期条目,提取对应信息并过滤冗余内容。

修改后的Selenium代码

from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# 点击展开营业时间按钮(保留原有逻辑)
span_xpath = "//span[contains(@class, 'x1q0g3np') and contains(@class, 'xohu8s8')]"
span_element = driver.find_element(By.XPATH, span_xpath)
button_inside_span = span_element.find_element(By.TAG_NAME, 'i')
button_inside_span.click()

# 等待营业时间容器加载完成
hours_container = WebDriverWait(driver, 10).until(
    EC.visibility_of_element_located((By.XPATH, '//div[@aria-label="Hours"]'))
)

# 定位每个星期的营业时间条目(适配Facebook页面结构)
day_entries = hours_container.find_elements(By.XPATH, './/div[contains(@class, "x193iq5w") and contains(@class, "xeuugli")]')

# 整理并过滤营业时间
formatted_hours = ["Hours"]
for entry in day_entries:
    # 提取星期名称
    day_name = entry.find_element(By.XPATH, './/span[contains(@class, "x193iq5w")]').text.strip()
    # 提取对应营业时间
    time_text = entry.find_element(By.XPATH, './/span[contains(@class, "x1lliihq")]').text.strip()
    # 过滤无关内容(热门时段、当前关闭提示等)
    if not ("Popular times" in time_text or "Closed now" in time_text):
        formatted_hours.append(f"{day_name} {time_text}")

# 输出格式化结果
for line in formatted_hours:
    print(line)

关键优化点

  1. 精准定位条目:不再直接取整个容器的文本,而是定位每个独立的星期条目元素,确保只提取对应星期的营业时间。
  2. 过滤冗余内容:通过关键词判断,排除"Popular times"(热门时段)、"Closed now"(当前关闭)等无关信息。
  3. 结构化输出:将提取的信息整理成你需要的「星期+时间」格式,首行保留"Hours"标题。

注意事项

Facebook的页面class可能会动态更新,如果上述xpath失效,可以通过以下方式调整:

  • 右键检查星期条目元素,找到更稳定的属性(比如包含星期文本的span,或固定的父容器结构)
  • 改用基于文本的定位,例如//span[text()="Monday"]/../span[2]来获取周一的营业时间

内容的提问来源于stack exchange,提问作者affadam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 22:25:20