如何从Facebook页面下拉菜单提取格式化的营业时间数据
解决Facebook页面营业时间爬取格式混乱问题
你的问题出在直接提取了整个aria-label="Hours"元素的文本,导致包含了热门时段、当前状态等无关内容。要获取规整的「星期+营业时间」格式,需要定位到每个单独的星期条目,提取对应信息并过滤冗余内容。
修改后的Selenium代码
from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # 点击展开营业时间按钮(保留原有逻辑) span_xpath = "//span[contains(@class, 'x1q0g3np') and contains(@class, 'xohu8s8')]" span_element = driver.find_element(By.XPATH, span_xpath) button_inside_span = span_element.find_element(By.TAG_NAME, 'i') button_inside_span.click() # 等待营业时间容器加载完成 hours_container = WebDriverWait(driver, 10).until( EC.visibility_of_element_located((By.XPATH, '//div[@aria-label="Hours"]')) ) # 定位每个星期的营业时间条目(适配Facebook页面结构) day_entries = hours_container.find_elements(By.XPATH, './/div[contains(@class, "x193iq5w") and contains(@class, "xeuugli")]') # 整理并过滤营业时间 formatted_hours = ["Hours"] for entry in day_entries: # 提取星期名称 day_name = entry.find_element(By.XPATH, './/span[contains(@class, "x193iq5w")]').text.strip() # 提取对应营业时间 time_text = entry.find_element(By.XPATH, './/span[contains(@class, "x1lliihq")]').text.strip() # 过滤无关内容(热门时段、当前关闭提示等) if not ("Popular times" in time_text or "Closed now" in time_text): formatted_hours.append(f"{day_name} {time_text}") # 输出格式化结果 for line in formatted_hours: print(line)
关键优化点
- 精准定位条目:不再直接取整个容器的文本,而是定位每个独立的星期条目元素,确保只提取对应星期的营业时间。
- 过滤冗余内容:通过关键词判断,排除"Popular times"(热门时段)、"Closed now"(当前关闭)等无关信息。
- 结构化输出:将提取的信息整理成你需要的「星期+时间」格式,首行保留"Hours"标题。
注意事项
Facebook的页面class可能会动态更新,如果上述xpath失效,可以通过以下方式调整:
- 右键检查星期条目元素,找到更稳定的属性(比如包含星期文本的span,或固定的父容器结构)
- 改用基于文本的定位,例如
//span[text()="Monday"]/../span[2]来获取周一的营业时间
内容的提问来源于stack exchange,提问作者affadam
相关产品推荐
相关产品推荐

