You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup解析HTML结构异常,如何正确获取<li>元素?

解决HTML解析结构异常并正确提取
  • 元素内容
  • 问题原因

    Python内置的html.parser对不规范HTML的容错处理能力有限,当目标页面原始HTML存在标签闭合不规范的情况时,会导致解析后出现<li>标签嵌套的异常结构,无法正常提取每个独立的<li>元素。

    解决方案

    1. 替换HTML解析器

    改用lxml或html5lib解析器,这两个工具对不规范HTML的兼容性更强,能正确还原预期的<ul>和<li>层级结构。

    首先安装对应的依赖库:

    # 安装lxml
    pip install lxml
    
    # 或者安装html5lib
    pip install html5lib
    

    2. 修改解析代码

    将原代码中的解析器参数替换为'lxml'或'html5lib':

    # 原代码
    soup = BeautifulSoup(response.content, 'html.parser')
    
    # 修改后(二选一)
    soup = BeautifulSoup(response.content, 'lxml')
    # 或者
    soup = BeautifulSoup(response.content, 'html5lib')
    

    3. 正确打印每个
  • 元素内容
  • 修改解析器后,即可正常遍历每个<li>元素,直接在循环中添加打印逻辑:

    for ul in holidays_elements:
        for index, li in enumerate(ul.find_all('li')):
            # 打印<li>元素的文本内容
            print(f"第{index+1}条:{li.get_text().strip()}")
            # 或者打印整个<li>元素的HTML结构
            print(f"第{index+1}个元素:{li}")
            
            date, name = li.get_text().strip().split(' - ', 1)
            if date not in holidays:
                holidays[date] = name
    

    修改后的完整代码

    from datetime import datetime
    import requests
    from bs4 import BeautifulSoup
    
    
    def is_holiday_or_weekend():
        current_year = datetime.now().year
        today = datetime.now().strftime('%Y-%m-%d')
    
        url = f"https://www.kalendorius.today/nedarbo-dienos/{current_year}"
    
        session = requests.Session()
    
        try:
            session.get(url)
            headers = {
                'User-Agent': 'Mozilla/5.0',
                'Accept': 'application/json',
            }
    
            response = session.get(url, headers=headers)
            response.raise_for_status()
    
            # 替换为lxml解析器
            soup = BeautifulSoup(response.content, 'lxml')
            
            holidays_elements = soup.find_all('ul', class_='calendar-items-list')
            holidays = {}
            for ul in holidays_elements:
                for index, li in enumerate(ul.find_all('li')):
                    # 打印每个<li>的文本内容
                    print(f"假期信息:{li.get_text().strip()}")
                    
                    date, name = li.get_text().strip().split(' - ', 1)
                    if date not in holidays:
                        holidays[date] = name
    
            if today in holidays or datetime.now().weekday() >= 5:
                return True
    
            return False
    
        except requests.RequestException as e:
            print(f"Error fetching holiday data: {e}")
            return None
    
    # Usage
    if is_holiday_or_weekend():
        print("Today is a holiday or weekend.")
    else:
        print("Today is a regular working day.")
    

    内容的提问来源于stack exchange,提问作者Dmitrij Holkin

    相关产品推荐
    方舟 Agent Plan

    超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

    最近更新时间:2026.07.02 21:17:13