You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup解析网站时获取空标签问题求助

问题解决:BeautifulSoup解析出空

标签的原因及方案

你的问题出在页面内容是JavaScript动态渲染的:用requests获取到的是服务器返回的初始静态HTML,而页面上显示的电话等内容是页面加载后通过JS动态插入的,所以BeautifulSoup解析初始HTML时自然拿不到实际文本。

解决方法

方法一:用Selenium模拟浏览器抓取

Selenium会模拟真实浏览器加载页面,等待JS执行完成后再获取渲染后的完整HTML,这样就能拿到目标标签的内容。

步骤:

  1. 安装Selenium:pip install selenium
  2. 下载对应浏览器的驱动(比如ChromeDriver,要和你的Chrome版本匹配)
  3. 使用以下代码:
from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from bs4 import BeautifulSoup as BS
import time

# 替换为你的ChromeDriver路径
service = Service('chromedriver.exe')
driver = webdriver.Chrome(service=service)

driver.get('https://naturasiberica.ru/our-shops/omsk/')
time.sleep(2)  # 等待JS加载完成,可根据实际情况调整时间

html = BS(driver.page_source, 'lxml')
phones = html.find_all('p', 'original-shops__phone')
for phone in phones:
    print(phone.text.strip())

driver.quit()

方法二:直接请求数据接口(更高效)

很多动态加载的网站会通过API接口获取数据,你可以通过浏览器抓包找到对应的接口:

  1. 打开浏览器开发者工具(F12),切换到「Network」标签页
  2. 刷新页面,筛选「XHR」或「Fetch」类型的请求,找到加载店铺数据的接口
  3. 直接用requests请求该接口,获取JSON格式的原始数据

示例代码(假设找到的接口URL如下):

import requests

api_url = 'https://naturasiberica.ru/api/shops/omsk'  # 替换为实际接口URL
response = requests.get(api_url)
shop_data = response.json()

# 从JSON数据中提取电话信息,具体字段根据接口返回结构调整
for shop in shop_data['data']:
    print(shop['phone'])

内容的提问来源于stack exchange,提问作者Ms.kitty

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 07:30:52