requests_html与BeautifulSoup无法渲染JavaScript页面的问题求助
解决requests_html未渲染JavaScript导致爬取模板变量的问题
问题描述
编写了结合requests_html与BeautifulSoup的爬虫代码,用于爬取https://www.trackingmore.com/track/en/的物流追踪信息,但代码未正确渲染JavaScript页面,爬取结果为{{info.Date}}、{{info.StatusDescription}}这类模板变量,而非实际物流数据。
原代码
from requests_html import HTMLSession from bs4 import BeautifulSoup session = HTMLSession() def track(num): url = f'https://www.trackingmore.com/track/en/{num}' headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 6.1; Win64; x64; rv:103.0) Gecko/20100101 Firefox/103.0'} r = session.post(url,headers=headers) # r.html.render(timeout=20) res = [] soup = BeautifulSoup(r.content,'lxml') st = soup.find('div',class_ ="track-status uk-flex") print(st.text) if st != 'Not Found': checkpoint = soup.find_all('div', class_="info-checkpoint") for i in checkpoint: date = i.find('div',class_='info-date').text.strip() desc = i.find('div',class_='info-desc').text.strip() res.append({ 'Date':date.replace('\xa0','')[:19], 'Description':desc.replace('\xa0','') }) return res else : return res
当前输出
[{'Date': '{{info.Date}} {{inf', 'Description': '{{info.StatusDescription}}'}, {'Date': '{{info.Date}} {{inf', 'Description': '{{info.StatusDescription}}'}]
解决方案
问题核心是未启用requests_html的JS渲染功能,同时存在请求方式错误和逻辑判断问题,修正如下:
修正后的代码
from requests_html import HTMLSession from bs4 import BeautifulSoup session = HTMLSession() def track(num): url = f'https://www.trackingmore.com/track/en/{num}' headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 6.1; Win64; x64; rv:103.0) Gecko/20100101 Firefox/103.0'} # 改用GET请求,该页面通过GET传递追踪号 r = session.get(url, headers=headers) # 启用JS渲染,等待页面加载完成 r.html.render(timeout=20) res = [] # 使用渲染后的HTML内容初始化BeautifulSoup soup = BeautifulSoup(r.html.html, 'lxml') st = soup.find('div', class_="track-status uk-flex") if st and 'Not Found' not in st.text.strip(): checkpoint = soup.find_all('div', class_="info-checkpoint") for i in checkpoint: date_elem = i.find('div', class_='info-date') desc_elem = i.find('div', class_='info-desc') # 避免元素不存在导致报错 if date_elem and desc_elem: date = date_elem.text.strip().replace('\xa0', '')[:19] desc = desc_elem.text.strip().replace('\xa0', '') res.append({'Date': date, 'Description': desc}) return res else: return res
关键修改说明
- 启用
r.html.render(timeout=20):这是requests_html渲染JavaScript页面的核心方法,会启动无头浏览器加载页面动态内容。 - 替换请求方式为
GET:该网站通过GET请求传递追踪号,原POST请求不符合页面逻辑。 - 使用
r.html.html作为BeautifulSoup输入:渲染后的动态内容存储在html属性中,而非原始的r.content。 - 修正"Not Found"判断逻辑:先判断
st是否存在,再检查文本内容是否包含"Not Found",避免对象与字符串直接比较的错误。 - 添加元素存在性判断:防止页面结构变化导致
find方法返回None,调用text时抛出异常。
内容的提问来源于stack exchange,提问作者nix20536
相关产品推荐
相关产品推荐

