爬取拉丁网站新闻出现字符乱码问题求助
解决拉丁字符爬取乱码问题
问题核心在于请求响应阶段的编码处理错误——你虽然设置了文件写入的UTF-8编码,但如果从网页获取内容时解码就出错,最终写入的文本自然会乱码。
修复步骤与代码调整
- 强制指定响应编码:目标网站实际采用UTF-8编码,但
requests可能自动推断编码错误,需手动指定解码规则。 - 使用解码后的文本解析:改用
request.text(已解码的字符串)而非request.content(字节流)传入BeautifulSoup,避免二次解码出错。
修改后的完整代码
import requests from bs4 import BeautifulSoup def xeberi_oxu(url): request = requests.get(url) # 强制指定响应编码为UTF-8,确保字符解码正确 request.encoding = 'utf-8' bs4 = BeautifulSoup(request.text, "html.parser") news_content = {} news_content['title'] = bs4.find("h1", class_='full-post-title').text news_content['content'] = bs4.find("div", class_='full-post-article').text with open(output_file, 'a', encoding='utf8') as file: file.write(f"Başlıq:{news_content['title']}\n") file.write(f"{news_content['content']}\n\n") url = "https://xebertv.info/" output_file = 'xeberler.txt' request = requests.get(url) request.encoding = 'utf-8' bs4 = BeautifulSoup(request.text, "html.parser") xeberler = bs4.find_all("div", class_="last-posts-list") for xeber in xeberler: links = xeber.find_all("a") for link in links: href = link.get("href") xeberi_oxu(href)
额外排查方案
如果仍有乱码,可先用chardet检测响应内容的真实编码:
import chardet result = chardet.detect(request.content) print(result['encoding'])
将检测出的编码值替换代码中的request.encoding = 'utf-8'即可。
内容的提问来源于stack exchange,提问作者TheFlee
相关产品推荐
相关产品推荐

