You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python正确抓取孟加拉语字符 解决乱码报错问题

Python抓取孟加拉语网页乱码解决方案

问题原因

  • requests库默认会基于响应头的charset字段推断网页编码,若站点未返回正确的charset声明,就会出现编码误判,导致孟加拉语这类非拉丁字符显示为乱码。
  • 本次抓取的站点实际编码为UTF-8,requests编码判断错误是乱码的核心诱因。

修复方案

有两种可用的修复方式,任选其一即可:

  1. 手动指定requests响应编码为UTF-8
  2. 直接将字节流传入BeautifulSoup,由BeautifulSoup自动识别编码

修复后代码

from bs4 import BeautifulSoup
import requests
import urllib3
import pprint
import json
from requests.models import DecodeError
from urllib3.util.url import Url

content = []

# for i in range(1, 25):
url = 'https://sattacademy.com/job-solution/view?cat_id=1&sub_cat_id=893'
res = requests.get(url)

# 修复代码二选一即可
# 方案1:手动指定编码
res.encoding = 'utf-8'
soup = BeautifulSoup(res.text, 'html.parser')

# 方案2:传字节流让BeautifulSoup自动识别,无需手动指定编码
# soup = BeautifulSoup(res.content, 'html.parser')

anchor = soup.findAll("a", attrs={"class": "nav-link"})

for ref in anchor:
    print(ref.text)

额外优化建议

  • 原有代码重复导入了两次requests,可删除多余的导入语句减少冗余。
  • 优先使用方案2,BeautifulSoup内置的编码检测能力适配性更强,遇到其他编码的小语种站点也不容易出现乱码问题。

内容的提问来源于stack exchange,提问作者Alif Hasan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 16:24:03