使用urlopen解码字节失败,正则匹配报错如何解决?
问题修复方案
错误根源
你调用re.search时参数顺序完全搞反了:re.search要求第一个参数是正则模式,第二个是要搜索的目标字符串,但你写成了search(page, "<body class=\"question-page")——把网页内容(page)当成正则表达式去解析,而网页内容里包含了无效的正则字符范围(比如[t-p],正则中字符范围必须是ASCII码从小到大,t的码值比p大,这种反向范围不合法),所以触发了正则解析错误。
修复方案
修正正则参数顺序
把正则模式放在第一个位置,网页内容放在第二个位置:import re from urllib.request import urlopen quest = input("question link: ").strip() page = urlopen(quest).read().decode() # 修正参数顺序 if re.search("<body class=\"question-page", page): # 你的后续逻辑 pass优化编码解码(可选但推荐)
直接用decode()可能会因网页编码不匹配导致乱码,建议从响应头获取正确编码:from urllib.request import urlopen quest = input("question link: ").strip() response = urlopen(quest) # 从响应头提取编码,默认用utf-8兜底 encoding = response.info().get_content_charset() or "utf-8" page = response.read().decode(encoding)用HTML解析库替代正则(推荐)
正则处理HTML容易出现各种问题,建议用BeautifulSoup这类专门的解析库:from urllib.request import urlopen from bs4 import BeautifulSoup quest = input("question link: ").strip() response = urlopen(quest) encoding = response.info().get_content_charset() or "utf-8" page = response.read().decode(encoding) soup = BeautifulSoup(page, "html.parser") # 直接查找body标签和对应class if soup.find("body", class_="question-page"): # 你的后续逻辑 pass
内容的提问来源于stack exchange,提问作者Wolf
相关产品推荐
相关产品推荐

