Python对URL执行二分查找识别字符'/'报错如何解决
报错根本原因
- 参数赋值错误:循环中
low = [i][0]、high = [i][-1]的写法得到的是完整URL字符串,不是索引值,把字符串传入二分查找做数值运算,直接触发类型错误。 - 函数逻辑混乱:二分查找函数的参数定义、递归调用的参数传递全是错的,原本要操作的单条URL和URL列表的逻辑完全混淆。
- 索引计算错误:用普通除法
/得到的是浮点数,不能作为字符串的索引取值,需要用整数除法//。 - 缺少递归边界:没有判断
low > high的终止条件,查找不存在的字符时会无限递归报错。
实现思路
首先明确二分查找的核心前提是操作的序列必须是有序的,针对URL的二分查找分两种常见场景:
- 在单个URL字符串中查找指定字符:Python字符默认按ASCII码比较,
/的ASCII码固定,直接即可比较,不需要额外做类型转换。 - 在URL列表中查找指定URL:需要先把URL列表按字典序排序,再对有序列表执行二分查找,Python原生支持URL字符串的字典序比较。
修正后可运行代码
场景1:单URL中查找是否存在/
url1 = "https://diversity.google" url2 = "https://www.aboutamazon.com/workplace/diversity-inclusion" url3 = "https://www.indeed.com/q-Diversity-jobs.html?vjk=ba073b4704d48c67" url4 = "https://careers.linkedin.com/diversity-and-inclusion" url5 = "https://github.com/about/diversity" url6 = "https://www.apple.com/diversity/" url7 ="https://www.samsung.com/us/about-us/diversity-and-inclusion/" url8 = "https://diversity.fb.com" url9 ="instagram:none" url10 = "https://careers.twitter.com/en/diversity.html" data = [url1,url2,url3,url4,url5,url6,url7,url8,url9,url10] # 修正后的二分查找:在target_str中查找x字符 def binary_search(target_str, low, high, x): # 递归边界:未找到字符 if high < low: return -1 # 整数除法得到合法整数索引 mid = (high + low) // 2 if target_str[mid] > x: return binary_search(target_str, low, mid - 1, x) elif target_str[mid] < x: return binary_search(target_str, mid + 1, high, x) else: # 找到字符返回对应索引 return mid x = '/' counterData = [] for url in data: low = 0 high = len(url) - 1 found_index = binary_search(url, low, high, x) # 返回值不等于-1则说明存在/ counter = 1 if found_index != -1 else 0 counterData.append(counter) print(counterData)
场景2:URL列表中查找指定URL
# 先对URL列表排序,满足二分查找的有序要求 sorted_data = sorted(data) def list_binary_search(sorted_list, low, high, target): if high < low: return -1 mid = (high + low) // 2 if sorted_list[mid] > target: return list_binary_search(sorted_list, low, mid -1, target) elif sorted_list[mid] < target: return list_binary_search(sorted_list, mid +1, high, target) else: return mid # 测试查找指定URL target_url = "https://diversity.fb.com" res = list_binary_search(sorted_data, 0, len(sorted_data)-1, target_url) print(f"目标URL索引为{res}" if res != -1 else "未找到目标URL")
注意事项
- 如果你只是要统计每个URL中
/的总个数,直接调用字符串内置方法url.count('/')即可,不需要用二分查找,二分查找只能判断字符是否存在、返回其中一个匹配位置,不能统计总数。 - 不要对无序列表/字符串执行二分查找,否则返回结果完全不可靠。
内容的提问来源于stack exchange,提问作者tquigg96
相关产品推荐
相关产品推荐

