如何使用Python从URL下载文本文件?解决字节格式问题
Python 下载文本文件的实现方法
方法一:用requests库直接获取并保存文本
requests库提供了更简洁的API处理HTTP请求,能直接获取解码后的文本内容:
import requests url = "目标文本文件的URL" # 发送GET请求获取响应 response = requests.get(url) # 自动识别文件真实编码,避免乱码 response.encoding = response.apparent_encoding # 将文本内容写入本地文件 with open("下载的文件.txt", "w", encoding="utf-8") as file: file.write(response.text)
response.text会自动把响应的字节数据解码为字符串,设置response.encoding = response.apparent_encoding能让requests根据内容自动推断正确编码,解决大部分乱码问题。
方法二:处理已获取的字节数据
如果你已经通过urlopen(url).read()拿到了字节形式的内容,只需用decode()方法转换为字符串后保存:
from urllib.request import urlopen url = "目标文本文件的URL" # 获取字节内容 byte_data = urlopen(url).read() # 用对应编码解码为字符串,比如utf-8、gbk等,根据文件实际编码调整 text_data = byte_data.decode("utf-8") # 写入本地文件 with open("下载的文件.txt", "w", encoding="utf-8") as file: file.write(text_data)
如果不确定文件编码,可以用chardet库自动检测:
import chardet from urllib.request import urlopen url = "目标文本文件的URL" byte_data = urlopen(url).read() # 检测编码 detected_encoding = chardet.detect(byte_data)["encoding"] text_data = byte_data.decode(detected_encoding) with open("下载的文件.txt", "w", encoding="utf-8") as file: file.write(text_data)
内容的提问来源于stack exchange,提问作者Yaver Javid
相关产品推荐
相关产品推荐

