调用profile()运行代码遇ValueError: math domain error求助
math domain error问题并完成性能对比图表 看起来你遇到的问题确实是math.log()传入了非正数导致的——对数函数的定义域只包含正数,当你的代码运行速度太快时,time.time() - start_time可能会返回0(甚至在部分系统上因为时钟精度问题出现极小负数),这时候就触发了ValueError。朋友那边能正常运行,大概率是他的系统时钟精度更高,或者CPU性能差异让时间差不会落到0值。
下面给你几个可行的解决思路,按从简单到严谨的顺序排列:
1. 给时间差加兜底的极小值
这是最快解决报错的方法,直接保证传入math.log()的数值是正数。你可以在调用方法后处理,或者修改方法的返回值:
比如修改append的代码:
time_method_in.append(math.log(max(method_in(a,b,c), 1e-9))) time_method_set_in.append(math.log(max(method_set_in(a,b,c), 1e-9))) time_method_bisect.append(math.log(max(method_bisect(a,b,c), 1e-9)))
这里1e-9是一个极小的正数,不会对最终的对数结果产生明显影响,却能完美避免0值问题。
2. 改用更高精度的计时函数
time.time()的精度在不同系统上表现不一,改用time.perf_counter()会更适合短时间的性能测量——它专门为性能测试设计,精度更高。
把所有方法里的time.time()替换成time.perf_counter():
# 比如method_in里的代码 start_time = time.perf_counter() # ... return time.perf_counter() - start_time
这样能更准确捕捉到短代码的运行时长,减少出现0值的概率。
3. 多次运行取平均(更严谨的性能测试)
单次运行的时间差误差极大,甚至会出现0值,这本身也会让你的性能对比结果不够准确。更好的做法是让每个方法重复运行多次,取平均时间:
修改方法示例:
def method_in(a,b,c): total_time = 0 runs = 5 # 可以根据N的大小调整,N越小可以增加运行次数 for _ in range(runs): # 每次运行前重置c的值,避免之前的运行结果干扰 c[:] = [0]*len(a) start_time = time.perf_counter() for i,x in enumerate(a): if x in b: c[i] = 1 total_time += time.perf_counter() - start_time return total_time / runs
这样不仅能避免0值问题,还能让你的性能数据更可靠,毕竟性能测试本来就需要多次采样来抵消偶然误差。
为什么log10和decimal没用?
不管是math.log()还是math.log10(),它们的定义域都是正数,只要传入0或者负数都会报错;decimal模块只是提升了数值精度,但如果原始时间差是0,decimal.Decimal(0).ln()同样会抛出错误——核心问题还是要保证传入对数函数的是正数,而不是换用不同的对数实现。
整合优化后的完整代码示例
这里把上面的优化点整合起来,你可以直接运行:
import random import bisect import matplotlib.pyplot as plt import math import time def method_in(a,b,c): total_time = 0 runs = 5 for _ in range(runs): c[:] = [0]*len(a) start_time = time.perf_counter() for i,x in enumerate(a): if x in b: c[i] = 1 total_time += time.perf_counter() - start_time return total_time / runs def method_set_in(a,b,c): total_time = 0 runs = 5 for _ in range(runs): c[:] = [0]*len(a) start_time = time.perf_counter() s = set(b) for i,x in enumerate(a): if x in s: c[i] = 1 total_time += time.perf_counter() - start_time return total_time / runs def method_bisect(a,b,c): total_time = 0 runs = 5 for _ in range(runs): c[:] = [0]*len(a) start_time = time.perf_counter() b_sorted = sorted(b) # 提前排序,避免每次循环重复排序 for i,x in enumerate(a): index = bisect.bisect_left(b_sorted,x) if index < len(b_sorted): if x == b_sorted[index]: c[i] = 1 total_time += time.perf_counter() - start_time return total_time / runs def profile(): time_method_in = [] time_method_set_in = [] time_method_bisect = [] Nls = [x for x in range(1000,20000,1000)] for N in Nls: a = list(range(0,N)) random.shuffle(a) b = list(range(0,N)) random.shuffle(b) c = [0]*len(a) # 加兜底值确保正数 time_method_in.append(math.log(max(method_in(a,b,c), 1e-9))) time_method_set_in.append(math.log(max(method_set_in(a,b,c), 1e-9))) time_method_bisect.append(math.log(max(method_bisect(a,b,c), 1e-9))) plt.plot(Nls,time_method_in,marker='o',color='r',linestyle='-',label='in') plt.plot(Nls,time_method_set_in,marker='o',color='b',linestyle='-',label='set') plt.plot(Nls,time_method_bisect,marker='o',color='g',linestyle='-',label='bisect') plt.xlabel('list size', fontsize=18) plt.ylabel('log(time)', fontsize=18) plt.legend(loc = 'upper left') plt.show() if __name__ == "__main__": profile()
另外注意我还优化了method_bisect里的排序逻辑——原来的代码每次循环都给b排序,这会额外增加不必要的耗时,现在改成提前排序好b_sorted,让测试更公平。
内容的提问来源于stack exchange,提问作者Smith Lo

