You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

爬取拉丁网站新闻出现字符乱码问题求助

解决拉丁字符爬取乱码问题

问题核心在于请求响应阶段的编码处理错误——你虽然设置了文件写入的UTF-8编码,但如果从网页获取内容时解码就出错,最终写入的文本自然会乱码。

修复步骤与代码调整

  1. 强制指定响应编码:目标网站实际采用UTF-8编码,但requests可能自动推断编码错误,需手动指定解码规则。
  2. 使用解码后的文本解析:改用request.text(已解码的字符串)而非request.content(字节流)传入BeautifulSoup,避免二次解码出错。

修改后的完整代码

import requests
from bs4 import BeautifulSoup

def xeberi_oxu(url):
    request = requests.get(url)
    # 强制指定响应编码为UTF-8,确保字符解码正确
    request.encoding = 'utf-8'
    bs4 = BeautifulSoup(request.text, "html.parser")
    news_content = {}
    news_content['title'] = bs4.find("h1", class_='full-post-title').text
    news_content['content'] = bs4.find("div", class_='full-post-article').text
    
    with open(output_file, 'a', encoding='utf8') as file:
        file.write(f"Başlıq:{news_content['title']}\n")
        file.write(f"{news_content['content']}\n\n")

url = "https://xebertv.info/"
output_file = 'xeberler.txt'
request = requests.get(url)
request.encoding = 'utf-8'
bs4 = BeautifulSoup(request.text, "html.parser")
xeberler = bs4.find_all("div", class_="last-posts-list")

for xeber in xeberler:
    links = xeber.find_all("a")
    for link in links:
        href = link.get("href")
        xeberi_oxu(href)

额外排查方案

如果仍有乱码,可先用chardet检测响应内容的真实编码:

import chardet
result = chardet.detect(request.content)
print(result['encoding'])

将检测出的编码值替换代码中的request.encoding = 'utf-8'即可。

内容的提问来源于stack exchange,提问作者TheFlee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 09:35:28