What Is Ollaya? Local Decision Models and How They Compare with Jev
Ollaya runs open decision models on your own hardware. Learn how it works, which models it supports, and when a local setup makes sense compared with Jev.

Suppose you want an app to sort incoming bug reports. You have three categories, a few examples, and a laptop. You probably do not want to spend the afternoon wiring up a different Python environment for every model you try.
Ollaya helps with that setup. It downloads and serves open decision models through a common local interface. You can try different models while keeping much of your application code the same.
The important distinction is that the runner and the model have different jobs. Ollaya manages execution; a model such as Laya or Winnow makes the judgment. That distinction explains why changing the model can change accuracy, memory use, and response time so much.
Here is how the project works, how it compares with Jev, and what is worth testing before you choose a local setup. Features and documentation were checked on October 2, 2026.
What is Ollaya?
Ollaya is an independent, open-source runtime for decision models. It provides a command line, a local server, and a desktop app. Its official GitHub repository describes a workflow built around downloading named models and serving typed decisions through a TypeSafe-compatible API.
A decision request contains some context, called the state, and questions about it. The state might be a bug report, an email, or a JSON record. The questions define what your application needs to know.
For example, given “The export button stopped working after yesterday's update,” you might ask:
- Which team owns this: product support, billing, or engineering?
- Does the report describe a broken feature?
- How much does the issue interrupt the user's work?
The output is a category, probability, or score that code can read. You still decide what happens next. A routing result can suggest a queue, but it does not fix the bug or contact the customer by itself.
Why people are looking at local decision models
The useful questions around Ollaya tend to be practical: can it run on a laptop, can existing Jev code call it, and will a small model make reliable choices?
An early developer discussion on Reddit describes using it behind a Python library for yes/no/unsure judgments. The author also reports model mistakes and slower first calls. That is one developer's experience, rather than a general performance result, but it highlights two real evaluation needs: handling uncertainty and measuring startup time.
Local processing also appeals to teams that want control over where their text is evaluated. Others simply want to try several open models without learning a new serving stack for each one. Those are good reasons to investigate the project. They do not tell you which model will work best on your documents.
The features that make Ollaya useful
One interface for several models
The project provides familiar commands for pulling, running, listing, and stopping models. Its registry points to the authors' weight files, with pinned revisions and checksum verification. The runner itself is Apache-2.0; model licenses need to be checked separately.
For a developer, the benefit is a shorter path from “I want to test this model” to “I have a result.” Keep the exact model tag in your experiment notes. A result labeled only “local AI” will be hard to reproduce a week later.
Typed questions your app can use
The API reference documents three question types:
| Type | What it gives you | Example question |
|---|---|---|
| Choice | A selected label and probabilities across options | Which team should handle this report? |
| Score | A value across an ordered set of described levels | How badly is the user blocked? |
| Noul | A probability for a yes/no judgment | Does the user report losing data? |
The wording matters. “Is this important?” is difficult to evaluate consistently. “Does the report say the user cannot complete a paid task?” gives you a clearer criterion and a result you can check against the original text.
Local execution with reusable configuration
According to the project FAQ, the server binds to the loopback address by default and evaluates inputs locally. Model downloads require network access. Your wider application still determines where inputs are stored or sent afterward.
You can also save recurring questions in presets or package a setup in a Modelfile. That is useful when the same issue categories need to stay consistent across a team's experiments. It does not, by itself, retrain the model.
Which models can Ollaya run?
The catalog includes both compact classifiers and larger decoder models. These options make different compromises, so start with a small shortlist.
| Model family | What it brings | A useful first test |
|---|---|---|
| Laya | English and multilingual decision checkpoints; a language-routing alias | Short messages on modest hardware |
| Winnow | Gemma-based decision models served from GGUF files | More demanding decisions when you have enough memory |
| Decider | Several Qwen-based sizes, including a vision variant | Compare model sizes on the same task |
| Kev | Decision models with a trained pointer head | Similar labels that are difficult to separate |
| NLI | Classifiers that judge whether text supports a hypothesis | Concrete statements such as “requests a replacement” |
| GLiClass | Classification with multiple candidate labels | Test your real category list together |
These are suggested experiments, not rankings. The official model catalog documents the available families and tags. Our Laya guide and Winnow guide provide more background on those models.
Choose an explicit tag when comparing results. The laya alias routes between English and multilingual checkpoints. If you already know the language, an explicit checkpoint makes it easier to see which model answered.
Ollaya vs Jev: what are you actually choosing?
Jev is TypeSafe's hosted decision model. Ollaya lets you run alternative models on hardware you manage. A shared request format makes comparison easier, but it does not make their judgments interchangeable.
| Consideration | Ollaya with an open model | Hosted Jev |
|---|---|---|
| Setup | Install the runner and download a model | Connect to a hosted service |
| Hardware | You provide memory and compute | Model compute is managed by the provider |
| Model choice | Choose from supported open checkpoints | Choose a supported Jev version |
| Data path | Local inference in the default setup | Send the request to the hosted provider |
| Operating cost | Hardware, electricity, and maintenance | Provider usage charges and integration costs |
| Performance | Depends on model, hardware, and workload | Depends on service conditions and workload |
| On this website | No local Ollaya connection is currently offered | Available in the Jev AI Playground |
TypeSafe's model reference describes Jev's versions, text input support, and limits. For a first experiment, a hosted Playground avoids the download and hardware setup. For a workflow that must run locally, you have a different requirement to satisfy.
A useful comparison keeps the inputs and category descriptions identical. If one model gets a full bug report and another gets a shortened summary, you are testing two workflows as well as two models. Write that down before drawing conclusions.
Is Ollaya the same as Ollama or Laya?
The similar names make this confusing. Laya is a model family. Ollaya is software that can run it. Ollama is a separate project.
The Ollaya FAQ explains that its command-line experience borrows ideas from Ollama and that the projects are independent. Similar commands do not mean their APIs, supported models, or configuration files are interchangeable. Check the endpoint your application actually calls.
It is also possible to use Laya without this runtime, through other implementations. Think of the runner as one way to deploy a model, rather than part of the model's identity. That keeps “Which model should I choose?” separate from “How should I serve it?”
How to use the local API
Once you have installed Ollaya for your operating system, a small first run can look like this:
ollaya run laya --preset triage "The download link in my receipt no longer works."
The official quickstart explains that run starts the server if needed, pulls the model on first use, and loads it. Expect setup time before treating the first response as a speed measurement.
For your own questions, the native endpoint is POST /api/decide. This Bash example assumes the local server is running, Laya has been downloaded, and server API-key enforcement is disabled:
curl http://localhost:11435/api/decide \
-H "Content-Type: application/json" \
-d '{
"model": "laya",
"state": "The download link in my receipt no longer works.",
"questions": {
"team": {
"type": "choice",
"instructions": "Which team should investigate this problem?",
"criteria": {
"billing": "Incorrect charges or payment disputes",
"product_support": "Accessing or using a purchased product",
"review": "Insufficient information to choose a team"
}
}
}
}'
Here, model selects the local model, state supplies the evidence, and questions names the judgments you want. Within a question, type selects the answer format, instructions states the task, and criteria defines the allowed labels. The team key identifies this answer in the response.
This request follows the documented schema; we have not run a local model benchmark for this article. Inspect the actual response rather than assuming a particular confidence value.
Connecting existing TypeSafe code
Ollaya also exposes /v1/systemone. Its compatibility guide documents using the TypeSafe Python SDK 0.7.1 with a local base URL and a local model name.
Set TYPESAFE_BASE_URL to http://localhost:11435 and select the model you pulled instead of leaving a Jev model name in the request. If the server uses OLLAYA_API_KEY, the client key must match. The guide also recommends bypassing system proxies for localhost requests.
API compatibility saves integration work. Still check error handling, input limits, and probability thresholds when changing backends. A threshold chosen for Jev should not be copied to Laya without testing.
Hardware and speed: look past one headline number
A CPU can be a reasonable starting point for compact models. Larger models require more memory, and acceleration depends on the model backend, platform, and installation method. Check the requirements for the specific tag you intend to run.
Treat published speed figures as measurements of a particular setup. The catalog's latency numbers use short inputs on an RTX 4090; they are not promises for an ordinary laptop. Its typed-decisions scores also come from a particular evaluation set with imperfect label agreement.
Run your own small comparison and record:
- The first request after loading the model.
- Several requests while the model is already loaded.
- Short and long versions of the same input.
- Mistakes that would send work to the wrong team.
The API documentation also says requests to the same loaded model queue. Sending a burst of requests may increase waiting time. Measure the full round trip your application experiences, not just the model's internal execution time.
Start with a Jev baseline in the Playground
Before setting up a local runner, it helps to know whether your decision question is clear. Open the Jev AI Playground and choose Your own case.
Paste the broken-download message into Text to evaluate, add a Choice question, and use the billing, product support, and review labels above. Select Run Jev to get a live result.
Now change the message to “I paid twice, and the download link is broken.” Which issue should take priority? If two people disagree, clarify the routing rule before blaming a model. Keep the revised question for your later local comparison.
The Playground currently runs Jev, not Ollaya or its local models. Our Jev examples offer more scenarios; distinguish their illustrative outputs from the live results you produce. For a broader shortlist, see our Jev alternative guide.
Frequently asked questions
Is Ollaya free and open source?
The runtime is open source under Apache-2.0. Models have their own licenses, and running them still uses your hardware and time. Local execution removes a hosted inference call, but it does not make operating costs disappear.
Can it run Jev itself locally?
The project serves open alternatives through a Jev-compatible interface. That compatibility does not provide TypeSafe's hosted Jev weights.
Can Ollaya handle images?
Supported vision models can. The current native API documents decider:2b-vision with one PNG image per request. Do not assume every model or the TypeSafe-compatible endpoint accepts images.
Is it worth trying?
Yes, if local deployment or comparing open checkpoints solves a real problem for you. Start with one task and a few cases you can judge yourself. A model that reliably sorts your own reports is more useful than a leaderboard position you cannot reproduce.
Explore Jev AI
See a Jev decision in context.
Open the Jev AI Playground, compare an example, and turn your own text into a focused question when the live option is available.
Open the Jev AI Playground


