使用Seleneitor提取数据后,如何仅保留Today和Tomorrow字段?
解决方案
方法一:处理已提取的文本字符串
如果你已经拿到完整的文本内容,可以通过字符串分割和列表推导式筛选出需要的内容:
# 将文本按换行符分割为列表 lines = texto_columnas.split('\n') # 筛选仅包含"Today"和"Tomorrow"的项 filtered_content = [line for line in lines if line in ("Today", "Tomorrow")] # 合并为字符串(按需使用) result = '\n'.join(filtered_content) print(result)
运行后会输出:
Today
Tomorrow
方法二:直接定位目标元素(更高效)
与其提取整个列表的文本再过滤,不如直接定位包含"Today"和"Tomorrow"的元素,减少后续处理步骤:
# 使用XPath直接匹配目标列表项 target_items = driver.find_elements(By.XPATH, '/html/body/div[5]/div[1]/div[4]/div/section[4]/section/div[1]/ul/li[text()="Today" or text()="Tomorrow"]') # 提取并拼接文本 filtered_text = '\n'.join([item.text for item in target_items]) print(filtered_text)
这种方法更稳定,避免了无关文本的干扰,也更符合爬虫的最佳实践。
内容的提问来源于stack exchange,提问作者JLL
相关产品推荐
相关产品推荐

