You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取网页时部分标签缺失的问题求助

解决爬虫无法获取完整HTML标签的常见方案
  • 动态内容渲染导致缺失:浏览器看到的部分标签是通过JavaScript动态加载生成的,requests/urllib只能获取服务器返回的静态HTML,无法执行JS。这种情况需要用浏览器自动化工具模拟加载过程:

    from selenium import webdriver
    from bs4 import BeautifulSoup
    from selenium.webdriver.support.ui import WebDriverWait
    from selenium.webdriver.support import expected_conditions as EC
    from selenium.webdriver.common.by import By
    
    driver = webdriver.Chrome()
    driver.get("https://example.com")
    # 显式等待目标元素加载完成,避免提前解析
    WebDriverWait(driver, 10).until(EC.presence_of_element_located((By.TAG_NAME, "目标标签名")))
    soup = BeautifulSoup(driver.page_source, 'html.parser')
    driver.quit()
    
  • 请求头不完整被服务器拦截:服务器会识别爬虫请求,返回简化版HTML。需要添加浏览器级别的请求头伪装:

    import requests
    from bs4 import BeautifulSoup
    
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
        'Accept-Language': 'zh-CN,zh;q=0.9',
        'Referer': 'https://example.com'
    }
    response = requests.get("https://example.com", headers=headers, verify=False)
    soup = BeautifulSoup(response.content, 'html.parser')
    
  • HTML解析器兼容性问题:默认的html.parser对不规范的HTML处理能力有限,换用更强大的解析器:

    # 使用lxml解析器(需先安装:pip install lxml)
    soup = BeautifulSoup(response.content, 'lxml')
    # 或使用html5lib解析器(需先安装:pip install html5lib)
    soup = BeautifulSoup(response.content, 'html5lib')
    
  • 反爬机制限制:部分网站需要携带有效Cookie、Session才能获取完整内容,可先通过浏览器登录获取Cookie,再加入请求中:

    cookies = {"cookie_name": "cookie_value"}
    response = requests.get("https://example.com", headers=headers, cookies=cookies, verify=False)
    

内容的提问来源于stack exchange,提问作者Amzzz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 15:42:22