使用requests爬取Glosbe时URL编码字符未被正确转换问题解决
爬取Glosbe德语句子时URL编码问题解决
我尝试从Glosbe.com爬取德语句子,请求URL包含UTF-8字符,但请求完成后网站未将编码后的字符转换回UTF-8。
预期的请求URL应为:
https://glosbe.com/de/hu/abkühlen
但网站实际接收的请求URL里,搜索词是编码后的格式:
https://glosbe.com/de/hu/abk%C3%BChlen/
使用的代码如下:
def beautifulSoapPrepare(sourceLang,destLang,phrase): headers = { 'User-Agent': 'My User Agent 1.0', 'From': 'youremail@domain.example' # This is another valid field } url="https://glosbe.com/"+sourceLang+"/"+destLang+"/"+urllib.parse.quote(phrase)+"/" r = requests.get(url, "lxml",headers=headers) soup = BeautifulSoup(r.content,features="lxml") return soup
解决方案
问题出在URL末尾的斜杠上,删除该斜杠后即可正常触发网站对编码字符的解码,搜索到目标单词abkühlen。
修改前代码:
url="https://glosbe.com/"+sourceLang+"/"+destLang+"/"+urllib.parse.quote(phrase)+"/"
修改后代码:
url="https://glosbe.com/"+sourceLang+"/"+destLang+"/"+urllib.parse.quote(phrase)
内容的提问来源于stack exchange,提问作者mrad2
相关产品推荐
相关产品推荐

