Python代码中URL含保留字报错(InstanceReservedWordNotAllowedInEntryKey)处理方案
解决InstanceReservedWordNotAllowedInEntryKey报错
问题根源
报错提示Entry Key包含保留字:,是因为你代码中使用了urllib.parse.unquote(url)——这个操作会把URL中已编码的特殊字符(比如%3A)还原成原始的:,而目标系统不允许Entry Key存在该保留字符。
解决方案
不要对原URL做解码操作,反而需要对URL中的特殊保留字符(比如锚点部分的:)进行URL编码,将其转换为%3A,同时保留URL本身的必要结构字符(如://、#)不被编码。
修改后的代码示例
方式一:精准拆分URL各部分编码(推荐)
import urllib.parse from urllib.parse import urlparse, urlunparse if not pd.isnull(row['resource']) and isinstance(row['resource'], str) and row['resource'].strip() != "": url = row['resource'] # 拆分URL结构,仅对路径和锚点中的特殊字符编码 parsed_url = urlparse(url) encoded_path = urllib.parse.quote(parsed_url.path, safe='/') # 保留路径中的/ encoded_fragment = urllib.parse.quote(parsed_url.fragment) # 编码锚点所有特殊字符 # 重新拼接合法URL encoded_url = urlunparse(( parsed_url.scheme, parsed_url.netloc, encoded_path, parsed_url.params, parsed_url.query, encoded_fragment )) url_updated = encoded_url if not (url.startswith("http://") or url.startswith("https://")): url_updated = "http://" + url_updated term.resources = [{ "displayName": url, "url": url_updated }]
方式二:快速整体编码(适合简单场景)
如果不需要精细拆分URL,直接使用quote并指定安全字符:
import urllib.parse if not pd.isnull(row['resource']) and isinstance(row['resource'], str) and row['resource'].strip() != "": url = row['resource'] # 编码URL,保留://#/这些必要结构字符,其他特殊字符自动编码 encoded_url = urllib.parse.quote(url, safe=':/#') url_updated = encoded_url if not (url.startswith("http://") or url.startswith("https://")): url_updated = "http://" + url_updated term.resources = [{ "displayName": url, "url": url_updated }]
说明
两种方式都会把URL中锚点部分的:转换为%3A,既不会破坏URL的可访问性,又能满足目标系统对Entry Key的字符限制要求。
内容的提问来源于stack exchange,提问作者aca1803
相关产品推荐
相关产品推荐

