使用llama_cpp_dart v0.2.2加载GGUF模型失败问题求助
GGUF模型加载失败问题(llama_cpp_dart v^0.2.2)
尝试使用llama_cpp_dart v^0.2.2运行GGUF模型时,模型无法加载。初始化函数initialize执行过程中,底层CPP代码调用load_llama_model_from_file(传入模型路径指针与modelparams)返回null,触发LlamaException。
使用的代码
Future<void> initialize() async { if (_state == ModelState.loading || _state == ModelState.ready) return; _set(ModelState.loading, p: 0.0); try { final nativeLibDir = await _getNativeLibDir(); debugPrint('Native lib dir: $nativeLibDir'); Llama.libraryPath = '$nativeLibDir/libllama.so'; const model = 'model.gguf'; final dir = await getApplicationSupportDirectory(); final filePath = '${dir.path}/$model'; final fileExists = await File(filePath).exists(); if (!fileExists) { final data = await rootBundle.load('assets/models/$model'); // 一次性同步写入,避免文件句柄"忙"状态 // 这样后续C++引擎读取时更安全 final file = File(filePath); await file.writeAsBytes( data.buffer.asUint8List(data.offsetInBytes, data.lengthInBytes), mode: FileMode.write, flush: true, // 强制操作系统完成写入后再继续 ); } _set(ModelState.loading, p: 0.90); final cmd = LlamaLoad( path: filePath, modelParams: ModelParams()..nGpuLayers = 0, contextParams: ContextParams() ..nCtx = 512 ..nThreads = 4, samplingParams: SamplerParams()..temp = 0.8, verbose: true, ); _parent = LlamaParent(cmd); _parent!.stream.listen( (token) => debugPrint('TOKEN: $token'), onError: (e) => debugPrint('ISOLATE ERROR: $e'), onDone: () => debugPrint('ISOLATE DONE'), cancelOnError: false, ); debugPrint('Calling init()...'); await _parent!.init().timeout(const Duration(minutes: 5), onTimeout: () { throw Exception('init() timed out'); }); _set(ModelState.ready, p: 1.0); } catch (e, st) { _errorMsg = e.toString(); debugPrint('TinyLlama init error: $e\n$st'); _set(ModelState.error); } } void _set(ModelState s, {double? p}) { _state = s; if (p != null) _loadProgress = p; notifyListeners(); } static const _channel = MethodChannel('native_lib_path'); Future<String> _getNativeLibDir() async { return await _channel.invokeMethod<String>('getPath') ?? ''; }
错误栈追踪信息
I/flutter (22081): Calling init()... I/Quality (22081): Skipped: false 12 cost 201.68422 refreshRate 16603314 bit true processName com.example.frontend D/VRI[MainActivity](22081): registerCallbacksForSync syncBuffer=false D/BLASTBufferQueue(22081): [VRI[MainActivity]#0](f:0,a:1) acquireNextBufferLocked size=1080x2400 mFrameNumber=1 applyTransaction=true mTimestamp=1129899983992346(auto) mPendingTransactions.size=0 graphicBufferId=94837172862991 transform=0 D/VRI[MainActivity](22081): Received frameCommittedCallback lastAttemptedDrawFrameNum=1 didProduceBuffer=true syncBuffer=false W/Parcel (22081): Expecting binder but got null! D/VRI[MainActivity](22081): debugCancelDraw cancelDraw=false,count = 247,android.view.ViewRootImpl@5f45cfc D/VRI[MainActivity](22081): draw finished. D/VRI[MainActivity](22081): onFocusEvent true I/flutter (22081): CPU : NEON = 1 | ARM_FMA = 1 | LLAMAFILE = 1 | REPACK = 1 | I/flutter (22081): modelPath: /data/user/0/com.example.frontend/files/model.gguf I/flutter (22081): libraryPath: /data/app/~~GJjJDHgHXAOuGkVJm8rXmA==/com.example.frontend-rxovsU9_2IHs-z_cK78-Nw==/lib/arm64/libllama.so I/flutter (22081): TinyLlama init error: LlamaException: Error loading model: LlamaException: Failed to initialize Llama (LlamaException: Could not load model at /data/user/0/com.example.frontend/files/model.gguf) I/flutter (22081): I/Quality (22081): Skipped: false 1 cost 16.622011 refreshRate 16605110 bit true processName com.example.frontend D/ProfileInstaller(22081): Installing profile for com.example.frontend W/OnBackInvokedCallback(22081): OnBackInvokedCallback is not enabled for the application. W/OnBackInvokedCallback(22081): Set 'android:enableOnBackInvokedCallback="true"' in the application manifest.
内容的提问来源于stack exchange,提问作者Kevin L
相关产品推荐
相关产品推荐

