What Is a Decision Model?
Generative models reason in unconstrained paragraphs. Decision models decide in typed, calibrated state transitions. Here is why the next frontier of AI is not generation, but decisions.

For three years, the AI industry has treated every computational problem as a text generation problem.
When an autonomous agent needs to select a tool, verify a financial transaction, triage a security vulnerability, adjust a robotic actuator, or determine whether an action requires human approval, the industry standard has been to prompt a 70-billion-parameter language model and wait for it to generate an open-ended stream of tokens. Developers then deploy brittle regex patterns, speculative JSON parsers, and retry loops to extract a single word or boolean flag from the response.
That approach is not merely computationally wasteful—spending hundreds of serial autoregressive decoding steps to yield an output that could be represented in two bits. It is fundamentally architecturally mismatched.
Generative language models were trained to maximize the likelihood of the next token across an infinite textual distribution. They are discursive, conversational, and structurally prone to hallucination under uncertainty. But 95% of the control flow in software systems does not need an essay.
It needs a decision.

Language Models Reason. Decision Models Decide.
To understand the difference, consider the underlying objective functions.
A Large Language Model (LLM) optimizes autoregressive token likelihood over an open sequence:
Loss_LLM = - Σ log P(w_t | w_<t)
It generates a sequence of tokens w_1, w_2, ..., w_n sampled from an unconstrained vocabulary of over 100,000 possibilities. The model cannot declare that it lacks the evidence to answer without generating natural language phrases that mimic modesty; it cannot output a native probability distribution over an operational state space; and its self-reported confidence numbers are famously uncalibrated.
A Decision Model optimizes a fundamentally different objective. Given an explicit evidence state E and a structured decision space D, a Decision Model computes a direct mapping to a bounded, typed output y:
y_hat = argmax_y P(y | E, D)
Equipped with a strictly calibrated posterior distribution P(y | E) and a first-class operational primitive for principled abstention (DEFER).
| Architectural Dimension | Generative Language Models (LLMs) | Decision Models (e.g., Hanzo Kai) |
|---|---|---|
| Primary Objective | Open-ended text synthesis & generation | Discrete, verifiable state transitions |
| Output Space | Unconstrained token sequences (100k+ vocab) | Bounded, strictly typed schemas (Choice, Score, Predicate, Defer) |
| Confidence & Calibration | Uncalibrated; confabulated softmax over tokens | Empirically calibrated probabilities (P = 0.90 implies 90% accuracy) |
| State Representation | Ephemeral context window of token strings | Structured multimodal evidence plane with cryptographic provenance |
| Execution Topology | Autoregressive serial decoding (O(N) latency) | Parallel graph resolution & Decision Diffusion (O(1) forward pass) |
| Uncertainty Handling | Hallucinates plausible completions | Principled Abstention (DEFER to policy or human reviewer) |
| Auditability | Non-deterministic verbal explanations | Immutable, replayable Decision Packages with cryptographic hash proofs |
When you separate reasoning from deciding, system architecture becomes orders of magnitude simpler, faster, and more reliable.
The Five Architectural Pillars of a Decision Model
A true Decision Model is not an LLM prompted with "Respond with valid JSON". It is a dedicated neural architecture designed from the ground up for bounded evaluation. Hanzo's open-weight decision model, Kai 1, is built upon five architectural pillars:
1. Bounded, Typed Outputs (Zero String Parsing)
In production software, every interface between systems must be strictly typed. Asking a language model to produce a decision and hoping it adheres to a regex schema introduces catastrophic tail-risk.
Decision models eliminate text parsing entirely. Every output is natively typed into one of four primitives:
Choice<T>: One selection from an explicit, enumerated candidate set, accompanied by the complete probability distribution over all alternatives (p_1, p_2, ..., p_kwhereΣ p_i = 1.0).Score[min, max]: A quantitative rating bounded within fixed operational bounds, complete with confidence intervals.Predicate: A pure boolean verification (trueorfalse) with calibrated certainty.Defer: An explicit refusal to guess when evidence is insufficient or conflicting, handing execution over to a predefined fallback policy or human-in-the-loop.
Because the output layer is constrained to the valid schema at the logit level, a decision model physically cannot produce an out-of-bounds result, an illegal enum value, or a malformed payload.
// Calling Hanzo Kai 1 via the unified SDK:
const decision = await kai.decide({
question: "deploy_canary_to_production",
schema: {
type: "choice",
options: ["promote", "rollback", "quarantine", "defer_to_sre"],
},
evidence: [
{ type: "p99_latency_ms", value: 142, threshold: 150 },
{ type: "error_rate_delta", value: "+0.002%", threshold: "+0.01%" },
{ type: "canary_health_score", value: 0.98 },
],
});
// The result is directly typed, validated, and calibrated:
console.log(decision.result); // "promote"
console.log(decision.probability); // 0.964
console.log(decision.entropy); // 0.12 (low uncertainty)
2. The Unified Multimodal Evidence Plane
Decisions in the real world rarely depend on text alone. In critical domains—such as sovereign finance, robotic manipulation, automated medical triage, and industrial telemetry—decisions require evaluating heterogeneous signals simultaneously.

A decision model does not serialize tabular data, sensor feeds, and radar arrays into clumsy Markdown tables for an LLM to read. Instead, Kai utilizes specialized modality encoders:
- Telemetry & Time Series: Vectorized rolling windows and frequency domain features.
- Visual & Thermal: Dense visual tokens from vision backbones.
- Structured State: Graph and relational entity projections from the Hanzo Datastore.
- Code & Specs: AST-level semantic embeddings and formal requirements matrices.
All modalities project into a unified evidence plane. Most importantly, evidence retains cryptographic provenance: every input tensor is hashed and fingerprinted. If a decision is audited six months later, the exact input state can be bit-for-bit reconstructed.
3. Parallel Graph Resolution & Decision Diffusion
The biggest bottleneck in modern autonomous agents is serial execution. An agent loop that requires 50 intermediate decisions typically calls an LLM 50 times in sequence. At 500ms to 2s per completion, the agent spends minutes simply traversing its own control flow.
Furthermore, error propagation in serial LLM chains is catastrophic. If each step has a 95% success rate, a 20-step chain succeeds only:
0.95 ^ 20 ≈ 35.8%

Decision models solve this via Decision Diffusion over directed acyclic graphs (DAGs).
Instead of evaluating decisions serially, the entire decision program is submitted at once. In a single forward pass:
- Pass 1 (Immediate Convergence): 80–90% of nodes with unambiguous evidence resolve immediately with high confidence (
P > 0.99). - Pass 2 (Conditional Diffusion): Dependent nodes receive the locked states of their parents, refining their probability distributions.
- Pass 3 (Constraint Satisfaction): Any remaining high-entropy or conflicting nodes are evaluated against formal boundary constraints.
A complex workflow involving 100 interdependent operational decisions resolves in milliseconds, not minutes.
4. Rigorous Calibration and Principled Abstention
If an AI model tells you it is 99% confident, how often is it actually right?
For off-the-shelf generative models, self-reported confidence is notoriously poorly correlated with truth. Models suffer from severe overconfidence on out-of-distribution inputs, happily hallucinating convincing justifications for erroneous outputs.
In industrial, aerospace, and financial applications, an overconfident wrong answer is fatal. A Decision Model must be empirically calibrated:
E[ Indicator(y_pred == y_true) | P(y_pred) = p ] = p
When Kai outputs a probability of 0.85, that decision is empirically correct exactly 85% of the time.
Uncalibrated LLM (Overconfident) Calibrated Decision Model (Kai)
1.0 | / 1.0 | /
| / | /
A 0.8 | _____/ A 0.8 | /
C | / C | /
C 0.6 | / C 0.6 | /
| / | /
0.4 | _____/ 0.4 | /
| / | /
0.2 | / 0.2 | /
| / | /
0.0 +---------------------- 0.0 +----------------------
0.0 0.2 0.4 0.6 0.8 1.0 0.0 0.2 0.4 0.6 0.8 1.0
CONFIDENCE CONFIDENCE
When entropy exceeds a parameterized risk tolerance τ, the model does not guess. It invokes Principled Abstention:
Entropy(P(y | E)) = - Σ p_i * log(p_i) > τ ==> DEFER
The system yields execution back to human specialists or deterministic rulebooks, providing an explicit sensitivity gradient identifying which piece of missing evidence would reduce entropy the most.
5. Replayable Decision Packages (The Audit Trail)
In regulated enterprise environments, "the AI generated this text" is not a valid legal defense or compliance posture.
Every decision evaluated by Kai produces an immutable, cryptographically sealed Decision Package:
{
"decision_id": "dec_8f0a2c91b4",
"program_version": "fraud_triage_v2.4",
"timestamp": "2026-09-30T19:00:00Z",
"model": {
"name": "kai-1-core",
"weights_sha256": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855"
},
"evidence_fingerprints": {
"transaction_context": "sha256:7c9e...",
"biometric_telemetry": "sha256:4a1d...",
"historical_risk_profile": "sha256:9f8b..."
},
"posterior_distribution": {
"approve": 0.012,
"step_up_mfa": 0.976,
"terminate_session": 0.009,
"defer_to_fraud_desk": 0.003
},
"selected_action": "step_up_mfa",
"entropy": 0.098,
"signature": "ed25519:3b7a8f..."
}
If an auditor or incident response engineer queries a transaction years later, they can replay the Decision Package against the exact model revision and verify that the output was mathematically deterministic, tamper-proof, and compliant with policy.
Where Decision Models Belong in Your Stack
Decision Models do not replace generative models. They relieve generative models of the jobs they were never designed to do.
In modern enterprise architectures, systems thrive when tasks are separated cleanly between generation and control:
┌──────────────────────────┐
│ Incoming Workload │
└─────────────┬────────────┘
│
▼
┌──────────────────────────┐
│ Enso Router │
│ (Task Intent & Triage) │
└──────┬────────────┬──────┘
│ │
Generative Tasks │ │ Control & Decision Tasks
(Code, Research, Content) │ │ (Routing, Approvals, Safety)
▼ ▼
┌───────────────┐ ┌───────────────┐
│ Zen Models │ │ Kai Models │
│ (LLM / Reason)│ │ (Decision Mod)│
└───────────────┘ └───────┬───────┘
│
▼
┌───────────────┐
│Typed Decision │
│ Package │
└───────────────┘
- Zen: Hanzo's frontier generative reasoning models. Use Zen when you need deep textual research, multi-step chain-of-thought proofs, creative synthesis, and document drafting.
- Kai: Hanzo's open-weight decision model family. Use Kai when you need deterministic routing, tool selection, safety thresholds, trade execution, triage, and agent loop control.
- Enso: The high-throughput meta-router that inspects incoming tasks at the microsecond layer, routing generative tasks to Zen and decision tasks to Kai.
Open Weights. Sovereign Infrastructure.
Reliable decision infrastructure cannot be locked inside a closed third-party API that changes behavior without notice. When a model governs production deployments, physical systems, or financial risk, you must have complete sovereignty over its weights, runtime, and data privacy.
Kai 1 is open weights. You can run it:
- On the Hanzo Sovereign Cloud: Scaled globally with zero-egress guarantees and sub-millisecond latencies.
- In Your Own VPC or On-Premises: Deployable via Docker or Kubernetes on NVIDIA GPUs, AMD ROCm, or high-density CPU nodes.
- At the Edge: Fully disconnected for robotics, aerospace, and sovereign defense environments.
Stop asking language models to write essays when your software just needs a decision.
Welcome to the era of Decision Intelligence.
Decisions, Not Completions
Run Kai through the Hanzo Cloud API today, or download the open weights to deploy calibrated decision models inside your sovereign infrastructure.
Read more

Kai 1: Decisions, Not Completions
Kai is Hanzo's decision model, served by the Hanzo API. Give it explicit state and a listed question, and it returns a typed, calibrated answer — a choice, a score, a yes/no, or a reasoned deferral.

Top Pain Points for Investment Firms That Hanzo AI Solves
Identify key operational challenges in private equity, venture capital, and asset management — and discover how Hanzo's sovereign AI cloud, frontier reasoning models, and agentic workflows streamline sourcing, diligence, and execution.
The router is the registry
Hanzo Cloud has no checked-in OpenAPI file. The live route table is the source, projected three ways, and a bijection test fails the build if any projection drifts. 1,362 operations across 963 paths, generated per process at request time.
