Kai 1: Decisions, Not Completions
Kai is Hanzo's open-weight decision model. Give it explicit state and a listed question, and it returns a typed, calibrated answer — a choice, a score, a yes/no, or a reasoned deferral.

Most steps in production software don't need a paragraph. They need an answer the software can act on: which tool to call, whether the evidence is enough, how severe a risk is, whether a human must sign off.
Today we're introducing Kai 1, Hanzo's open-weight decision model. Give Kai the facts of a situation and a question with listed options. It returns a typed answer (a choice, a score, a yes/no, or a calibrated "not sure, defer") with a probability for every option.
Kai 1 is live on the Hanzo API, and its weights and code are open.
Language models reason. Kai decides.
When you ask a language model to make a call, you get text back. You have to parse it, trust a confidence number the model made up, and hope the output stays in bounds.
Kai is built for the steps where that isn't good enough. Its answers are bounded, its probabilities are calibrated, and there's nothing to parse. Every Kai answer is one of four types:
- Choice: one option from a listed set, with a probability for each
- Score: a rating against levels you define
- Predicate: a yes/no with its probability
- Defer: Kai isn't confident enough and hands the call to a person or a rule
Evidence stays attached
A decision shouldn't care whether its evidence started as text, an image or a sensor reading. Kai is designed to take mixed evidence and keep its origin attached to every answer.
A number without provenance is just a number.
Many decisions, one pass
Agent loops ask one question at a time. Kai is designed to run a whole decision program at once: one shared state, many decisions, with dependencies spelled out. Confident answers are locked in first, and only the uncertain ones get more work.
The goal is that a program with thousands of decisions shouldn't need thousands of serial model calls.
Every decision can be replayed
Each Kai decision records the program version, evidence fingerprints, model revision, full probability distribution, approvals and execution trace. When someone asks why the system did something, you can show them.
Where Kai fits
- Agents: model and tool selection, when to continue or stop, when to ask a human
- Business workflows: lead scoring, intent detection, next best action, escalation
- Physical systems: vehicles, robotics, manufacturing, energy and infrastructure, where decisions run on sensor data
Part of the Hanzo stack
Generation and decision are different jobs. Zen6 is our family of open-weight generation models. Kai 1 makes bounded, calibrated decisions. Enso routes each task to the one it needs.
Run it anywhere
Kai runs where your evidence lives. Call it through the Hanzo API, or take the open weights and run it in your own cloud, on Kubernetes, on-premises or fully disconnected.
What's next
We're working on seven open research questions, and we'll publish results only once they can be reproduced:
- Decisions over extremely large option sets
- Multilingual decisions
- Shared-state decoding
- Decision diffusion (refining only uncertain answers)
- Multimodal evidence
- Calibration and knowing when to abstain
- Consistency across connected decisions
Benchmark results for Kai aren't published yet. They'll go out as they pass reproducible release gates.
Stop asking one model to do everything
Call Kai through the Hanzo API today, or download the weights and run it where your evidence lives.
Read more

Top Pain Points for Investment Firms That Hanzo AI Solves
Identify key operational challenges in private equity, venture capital, and asset management — and discover how Hanzo's sovereign AI cloud, frontier reasoning models, and agentic workflows streamline sourcing, diligence, and execution.
The router is the registry
Hanzo Cloud has no checked-in OpenAPI file. The live route table is the source, projected three ways, and a bijection test fails the build if any projection drifts. 1,362 operations across 963 paths, generated per process at request time.
Half a Million RPS on a GB10: The Network Was the Bottleneck, Not the Box
We set out to see how fast our cloud data plane could go on an NVIDIA GB10. The answer taught us more about NIC receive queues than about our own code — and it more than doubled our throughput on the same wire.
