You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python bs4爬取TakeAlot网站出现乱码,请求排查原因

问题描述

使用Python的bs4库爬取TakeAlot网站时,仅返回少量HTML代码,其余内容为乱码,无法正常解析。使用的代码如下:

from bs4 import BeautifulSoup
import requests

htmlSource = (requests.get(
    "https://www.takealot.com/?gclid=CjwKCAjwzNOaBhAcEiwAD7Tb6AhZbkyxR6ewJSUswC-GGcilxY3D10zvFd4repwE3SGZDbVn7U6q4RoC5cwQAvD_BwE&gclsrc=aw.ds"
)).text
soup = BeautifulSoup(htmlSource, "html.parser")
print(soup)
问题排查与解决

核心原因

这种情况要么是网站反爬机制拦截了你的请求,要么是请求未正确处理网站返回的压缩编码,导致内容乱码。

解决步骤

  1. 模拟浏览器请求头
    网站会通过User-Agent识别请求来源,默认的requests请求头会被标记为爬虫。添加浏览器的User-Agent、Accept等字段,让请求更像正常用户访问。

  2. 自动处理压缩编码
    很多网站会返回gzip压缩的内容,直接取.text可能无法正确解码,需要让requests自动处理解压。

修改后的代码示例:

from bs4 import BeautifulSoup
import requests

headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36",
    "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8",
    "Accept-Encoding": "gzip, deflate, br",
    "Accept-Language": "en-US,en;q=0.5"
}

# 启用自动解压
response = requests.get(
    "https://www.takealot.com/?gclid=CjwKCAjwzNOaBhAcEiwAD7Tb6AhZbkyxR6ewJSUswC-GGcilxY3D10zvFd4repwE3SGZDbVn7U6q4RoC5cwQAvD_BwE&gclsrc=aw.ds",
    headers=headers,
    allow_redirects=True
)
# 确保requests自动解码压缩内容
response.encoding = response.apparent_encoding

soup = BeautifulSoup(response.text, "html.parser")
print(soup.prettify())
  1. 动态渲染页面的备选方案
    如果上述方法还是不行,说明网站内容是通过JavaScript动态加载的,requests无法获取渲染后的HTML。这时需要用selenium或playwright模拟真实浏览器加载页面:

以selenium为例的代码示例:

from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.chrome.options import Options

options = Options()
options.add_argument("--headless=new")  # 无头模式,不显示浏览器窗口
options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36")

driver = webdriver.Chrome(options=options)
driver.get("https://www.takealot.com/?gclid=CjwKCAjwzNOaBhAcEiwAD7Tb6AhZbkyxR6ewJSUswC-GGcilxY3D10zvFd4repwE3SGZDbVn7U6q4RoC5cwQAvD_BwE&gclsrc=aw.ds")

soup = BeautifulSoup(driver.page_source, "html.parser")
print(soup.prettify())

driver.quit()

注意事项

  • 爬取网站前务必查看其robots.txt文件,确认允许爬取的内容,避免违反网站规则。
  • 不要频繁发送请求,可添加适当的延迟,防止被封禁IP。

内容的提问来源于stack exchange,提问作者Aaron Vegoda

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 01:45:36