My writing assistant could talk to Gemini. The missing piece was using it without a connection.
That was the reason I started llamadart: I wanted an offline mode for flights, unreliable internet, and environments where outbound API access was unavailable. In my first article, I focused on the work needed to get there: native binaries, Dart build hooks, and model chat templates.
Getting a response from a local model is a satisfying milestone. It also leaves you with a surprisingly long list of application problems.
Where does the model come from? What happens if someone cancels its download? Where does the conversation live? Can they switch models without losing the one that already works?
Those questions are a useful way to explain how llamadart has changed. The project now does more, but the interesting part of 0.11 is how much of the ordinary integration work fits into a smaller, clearer API.
Let's use a simple writing-assistant flow to walk through it: download a model, ask it to rewrite a paragraph, then ask for a shorter version. This is an illustrative app flow, not a claim about a particular model's writing quality.
Before the first answer
Imagine adding an “Offline mode” button to that application. Pressing it cannot immediately make a large model appear on the device.
The app needs to explain the download, show progress, and recover when the user closes the screen. Later launches should reuse the file rather than download it again. These are the details a user encounters before they ever see a generated word.
The native packaging work from the first article still helps here. App developers do not need a local C++ toolchain for the common setup; the package's build hook resolves the native runtime assets. Model files are a separate concern.
In 0.11, the model-loading entrypoint brings those files together. A LlamaModel describes the model and an optional multimodal projector. Its ModelSource can be a local path, a URL, or a Hugging Face reference.
The change is small enough to show directly. With source and params already defined, the old setup looked like this:
// 0.10
final engine = LlamaEngine(LlamaBackend());
try {
await engine.loadModelSource(source, modelParams: params);
// Use the engine.
} finally {
await engine.dispose();
}
Now the ready engine comes back from the load:
// 0.11
final engine = await LlamaEngine.load(
LlamaModel(source),
params: params,
);
try {
// Use the engine.
} finally {
await engine.dispose();
}
The important difference is where ownership begins. If the new load throws, the engine and backend it created are already disposed. If it succeeds, the caller owns the ready engine.
Progress and cancellation are still things the app needs to present. The loader exposes them through onProgress and download options, so they can belong to the same operation as loading the model. The lifecycle guide covers that operation in detail.
From a prompt to a conversation
Once the model is ready, our writing assistant can ask for a rewrite. Then comes the follow-up: “Make it shorter.”
That short instruction depends on the earlier exchange. The app needs to keep the paragraph and the first answer in the conversation.
ChatSession handles that history. It already existed before 0.11; the new convenience methods make the common case easier to read. Send a message with session.send, then get the completed response through reply.text.
Here is a complete native Dart starting point:
import 'package:llamadart/llamadart.dart';
Future<void> main() async {
final engine = await LlamaEngine.load(
LlamaModel(
ModelSource.parse(
'hf://unsloth/SmolLM2-135M-Instruct-GGUF/'
'SmolLM2-135M-Instruct-Q2_K.gguf',
),
),
params: const ModelParams(contextSize: 1024, gpuLayers: 0),
);
try {
final session = ChatSession(
engine,
systemPrompt: 'You help rewrite text concisely.',
);
final reply = await session.send(
'Rewrite this: We are writing to let you know '
'that the meeting will start at nine.',
);
print(reply.text);
final shorter = await session.send('Make it shorter.');
print(shorter.text);
} finally {
await engine.dispose();
}
}
The tiny model keeps this example approachable. Choosing a model that follows your instructions well is a separate step, and you should evaluate it with your own inputs.
For a live chat UI, session.create still streams the response. Reading its text has changed from chunk.choices.first.delta.content to chunk.text. That is a small improvement you notice every time you connect the stream to a widget.
There is a limit to conversation memory. ChatSession trims older turns to fit the context budget while keeping the system prompt. If your app already owns and edits transcripts, engine.create lets you supply the full message list yourself.
These examples follow the tagged 0.10 and 0.11 documentation.
The model picker changes the problem
Now imagine the writing assistant has a model picker. Someone wants to try a different model while the current one is available.
An app could unload the current model first and then start downloading its replacement. On a slow connection, that leaves the user waiting without either model ready.
On native targets, setModel keeps the current model available while the replacement files resolve or download. If that stage fails or is cancelled, the old model stays loaded. Only after the files are ready does replacement begin.
This has a boundary worth understanding: a failure during the later load can leave nothing loaded. Web runtimes also unload before fetching the replacement. The API gives the app a defined lifecycle, rather than a guarantee that every switch succeeds.
The same ownership question appears when leaving the chat screen. A Flutter engine should live in a service, provider, or state object, not inside build(). Dispose it when that owner is finished. Desktop app exit needs its own cleanup path where documented.
These choices are less visible than a new model demo, but they determine whether the feature behaves sensibly when someone uses the rest of the app.
More runtimes without pretending they are identical
The project has also grown beyond its original llama.cpp path.
Version 0.7 added LiteRT-LM support. Version 0.8 separated Apple runtime companions from the core package, keeping pure Dart consumers free of a Flutter SDK constraint. Version 0.10 added preview image generation through an opt-in stable-diffusion.cpp runtime.
That creates more possibilities, along with more differences the application must respect.
For our writing assistant, a useful question is whether to enable an image attachment button. The answer should come from the loaded model and runtime. In 0.11, engine.runtime identifies that runtime, and await engine.capabilities reports supported operations.
Device selection follows the same principle. ModelParams.device provides a shared CPU, GPU, or NPU selection. An explicit request runs there or fails as unsupported; auto preserves the runtime's default.
There are still platform limits. Android llama.cpp Vulkan is experimental and device-dependent. Web support is experimental. Automatic tool loops reject the pinned LiteRT-LM runtimes because they cannot reliably report token-limit truncation. The support matrix is the place to check a deployment.
The payoff for app code is a common way to ask what is available, and a clear failure when a requested operation cannot be supported.
Try one useful offline feature
To run the example, add llamadart: ^0.11.0 and run dart pub get or flutter pub get. The package requires Dart 3.10.7 or newer, and Flutter 3.38.0 or newer for Flutter apps.
Flutter iOS and macOS apps using the GGUF Apple companion should pair it with llamadart_llama_cpp_flutter: ^0.0.21. Other runtimes have their own companions and requirements; follow the installation guide.
The first setup needs a connection for missing runtime assets and model files. Native model downloads are cached for reuse. Once the required assets are present, inference can run locally; any remote tools or cloud services you add still need their own connections.
For an existing app, read the migration guide. Older model-loading methods remain deprecated until 1.0, but other API and lifecycle changes need attention.
Start with one feature: a rewrite, a summary, or a conversation. Get it working, then try cancelling, switching models, and leaving the screen. Those interactions tell you much more about the integration than a successful first prompt.
The Dart and Flutter examples are there to build on. I would welcome feedback about the parts that still make that journey harder than it needs to be.
