You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

requests_html与BeautifulSoup无法渲染JavaScript页面的问题求助

解决requests_html未渲染JavaScript导致爬取模板变量的问题

问题描述

编写了结合requests_html与BeautifulSoup的爬虫代码,用于爬取https://www.trackingmore.com/track/en/的物流追踪信息,但代码未正确渲染JavaScript页面,爬取结果为{{info.Date}}、{{info.StatusDescription}}这类模板变量,而非实际物流数据。

原代码

from requests_html import HTMLSession
from bs4 import BeautifulSoup

session = HTMLSession()


def track(num):
    url = f'https://www.trackingmore.com/track/en/{num}'
    headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 6.1; Win64; x64; rv:103.0) Gecko/20100101 Firefox/103.0'}
    r = session.post(url,headers=headers)
    # r.html.render(timeout=20)
    res = []
    soup =  BeautifulSoup(r.content,'lxml')
    st = soup.find('div',class_ ="track-status uk-flex")
    print(st.text)
    if st != 'Not Found':
        checkpoint = soup.find_all('div', class_="info-checkpoint")
        for i in checkpoint:
            date = i.find('div',class_='info-date').text.strip()
            desc = i.find('div',class_='info-desc').text.strip()
            res.append({
                    'Date':date.replace('\xa0','')[:19],
                    'Description':desc.replace('\xa0','')
                    })
        return res
    else :
        return res

当前输出

[{'Date': '{{info.Date}} {{inf', 'Description': '{{info.StatusDescription}}'}, {'Date': '{{info.Date}} {{inf', 'Description': '{{info.StatusDescription}}'}]

解决方案

问题核心是未启用requests_html的JS渲染功能,同时存在请求方式错误和逻辑判断问题,修正如下:

修正后的代码

from requests_html import HTMLSession
from bs4 import BeautifulSoup

session = HTMLSession()

def track(num):
    url = f'https://www.trackingmore.com/track/en/{num}'
    headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 6.1; Win64; x64; rv:103.0) Gecko/20100101 Firefox/103.0'}
    # 改用GET请求,该页面通过GET传递追踪号
    r = session.get(url, headers=headers)
    # 启用JS渲染,等待页面加载完成
    r.html.render(timeout=20)
    res = []
    # 使用渲染后的HTML内容初始化BeautifulSoup
    soup = BeautifulSoup(r.html.html, 'lxml')
    st = soup.find('div', class_="track-status uk-flex")
    
    if st and 'Not Found' not in st.text.strip():
        checkpoint = soup.find_all('div', class_="info-checkpoint")
        for i in checkpoint:
            date_elem = i.find('div', class_='info-date')
            desc_elem = i.find('div', class_='info-desc')
            # 避免元素不存在导致报错
            if date_elem and desc_elem:
                date = date_elem.text.strip().replace('\xa0', '')[:19]
                desc = desc_elem.text.strip().replace('\xa0', '')
                res.append({'Date': date, 'Description': desc})
        return res
    else:
        return res

关键修改说明

  • 启用r.html.render(timeout=20):这是requests_html渲染JavaScript页面的核心方法,会启动无头浏览器加载页面动态内容。
  • 替换请求方式为GET:该网站通过GET请求传递追踪号,原POST请求不符合页面逻辑。
  • 使用r.html.html作为BeautifulSoup输入:渲染后的动态内容存储在html属性中,而非原始的r.content。
  • 修正"Not Found"判断逻辑:先判断st是否存在,再检查文本内容是否包含"Not Found",避免对象与字符串直接比较的错误。
  • 添加元素存在性判断:防止页面结构变化导致find方法返回None,调用text时抛出异常。

内容的提问来源于stack exchange,提问作者nix20536

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 12:15:41