What Is a Decision Model?

Generative models reason in unconstrained paragraphs. Decision models decide in typed, calibrated state transitions. Here is why the next frontier of AI is not generation, but decisions.

What Is a Decision Model?

For three years, the AI industry has treated every computational problem as a text generation problem.

When an autonomous agent needs to select a tool, verify a financial transaction, triage a security vulnerability, adjust a robotic actuator, or determine whether an action requires human approval, the industry standard has been to prompt a 70-billion-parameter language model and wait for it to generate an open-ended stream of tokens. Developers then deploy brittle regex patterns, speculative JSON parsers, and retry loops to extract a single word or boolean flag from the response.

That approach is not merely computationally wasteful—spending hundreds of serial autoregressive decoding steps to yield an output that could be represented in two bits. It is fundamentally architecturally mismatched.

Generative language models were trained to maximize the likelihood of the next token across an infinite textual distribution. They are discursive, conversational, and structurally prone to hallucination under uncertainty. But 95% of the control flow in software systems does not need an essay.

It needs a decision.

High-dimensional manifold visualizing deterministic, calibrated decision states bounded by mathematical constraints.
Figure 1 · Bounded state transitions over unconstrained token sequences. Decision models map multimodal evidence directly into calibrated action distributions.

Language Models Reason. Decision Models Decide.

To understand the difference, consider the underlying objective functions.

A Large Language Model (LLM) optimizes autoregressive token likelihood over an open sequence:

Loss_LLM = - Σ log P(w_t | w_<t)

It generates a sequence of tokens w_1, w_2, ..., w_n sampled from an unconstrained vocabulary of over 100,000 possibilities. The model cannot declare that it lacks the evidence to answer without generating natural language phrases that mimic modesty; it cannot output a native probability distribution over an operational state space; and its self-reported confidence numbers are famously uncalibrated.

A Decision Model optimizes a fundamentally different objective. Given an explicit evidence state E and a structured decision space D, a Decision Model computes a direct mapping to a bounded, typed output y:

y_hat = argmax_y P(y | E, D)

Equipped with a strictly calibrated posterior distribution P(y | E) and a first-class operational primitive for principled abstention (DEFER).

Architectural DimensionGenerative Language Models (LLMs)Decision Models (e.g., Hanzo Kai)
Primary ObjectiveOpen-ended text synthesis & generationDiscrete, verifiable state transitions
Output SpaceUnconstrained token sequences (100k+ vocab)Bounded, strictly typed schemas (Choice, Score, Predicate, Defer)
Confidence & CalibrationUncalibrated; confabulated softmax over tokensEmpirically calibrated probabilities (P = 0.90 implies 90% accuracy)
State RepresentationEphemeral context window of token stringsStructured multimodal evidence plane with cryptographic provenance
Execution TopologyAutoregressive serial decoding (O(N) latency)Parallel graph resolution & Decision Diffusion (O(1) forward pass)
Uncertainty HandlingHallucinates plausible completionsPrincipled Abstention (DEFER to policy or human reviewer)
AuditabilityNon-deterministic verbal explanationsImmutable, replayable Decision Packages with cryptographic hash proofs

When you separate reasoning from deciding, system architecture becomes orders of magnitude simpler, faster, and more reliable.


The Five Architectural Pillars of a Decision Model

A true Decision Model is not an LLM prompted with "Respond with valid JSON". It is a dedicated neural architecture designed from the ground up for bounded evaluation. Hanzo's open-weight decision model, Kai 1, is built upon five architectural pillars:

1. Bounded, Typed Outputs (Zero String Parsing)

In production software, every interface between systems must be strictly typed. Asking a language model to produce a decision and hoping it adheres to a regex schema introduces catastrophic tail-risk.

Decision models eliminate text parsing entirely. Every output is natively typed into one of four primitives:

  • Choice<T>: One selection from an explicit, enumerated candidate set, accompanied by the complete probability distribution over all alternatives (p_1, p_2, ..., p_k where Σ p_i = 1.0).
  • Score[min, max]: A quantitative rating bounded within fixed operational bounds, complete with confidence intervals.
  • Predicate: A pure boolean verification (true or false) with calibrated certainty.
  • Defer: An explicit refusal to guess when evidence is insufficient or conflicting, handing execution over to a predefined fallback policy or human-in-the-loop.

Because the output layer is constrained to the valid schema at the logit level, a decision model physically cannot produce an out-of-bounds result, an illegal enum value, or a malformed payload.

// Calling Hanzo Kai 1 via the unified SDK:
const decision = await kai.decide({
  question: "deploy_canary_to_production",
  schema: {
    type: "choice",
    options: ["promote", "rollback", "quarantine", "defer_to_sre"],
  },
  evidence: [
    { type: "p99_latency_ms", value: 142, threshold: 150 },
    { type: "error_rate_delta", value: "+0.002%", threshold: "+0.01%" },
    { type: "canary_health_score", value: 0.98 },
  ],
});

// The result is directly typed, validated, and calibrated:
console.log(decision.result);      // "promote"
console.log(decision.probability); // 0.964
console.log(decision.entropy);     // 0.12 (low uncertainty)

2. The Unified Multimodal Evidence Plane

Decisions in the real world rarely depend on text alone. In critical domains—such as sovereign finance, robotic manipulation, automated medical triage, and industrial telemetry—decisions require evaluating heterogeneous signals simultaneously.

Heterogeneous evidence streams—code, telemetry, documents, sensors—encoded into a unified high-dimensional evidence core with cryptographic provenance.
Figure 2 · The Multimodal Evidence Plane. Heterogeneous inputs are encoded into a shared evidence representation where every input retains cryptographic SHA-256 provenance.

A decision model does not serialize tabular data, sensor feeds, and radar arrays into clumsy Markdown tables for an LLM to read. Instead, Kai utilizes specialized modality encoders:

  • Telemetry & Time Series: Vectorized rolling windows and frequency domain features.
  • Visual & Thermal: Dense visual tokens from vision backbones.
  • Structured State: Graph and relational entity projections from the Hanzo Datastore.
  • Code & Specs: AST-level semantic embeddings and formal requirements matrices.

All modalities project into a unified evidence plane. Most importantly, evidence retains cryptographic provenance: every input tensor is hashed and fingerprinted. If a decision is audited six months later, the exact input state can be bit-for-bit reconstructed.


3. Parallel Graph Resolution & Decision Diffusion

The biggest bottleneck in modern autonomous agents is serial execution. An agent loop that requires 50 intermediate decisions typically calls an LLM 50 times in sequence. At 500ms to 2s per completion, the agent spends minutes simply traversing its own control flow.

Furthermore, error propagation in serial LLM chains is catastrophic. If each step has a 95% success rate, a 20-step chain succeeds only:

0.95 ^ 20 ≈ 35.8%
A dense interconnected decision graph resolving across parallel diffusion passes, locking high-confidence nodes instantly and refining high-entropy edges.
Figure 3 · Decision Diffusion. Instead of dozens of sequential LLM round-trips, a decision DAG settles in parallel passes, refining only high-entropy nodes.

Decision models solve this via Decision Diffusion over directed acyclic graphs (DAGs).

Instead of evaluating decisions serially, the entire decision program is submitted at once. In a single forward pass:

  1. Pass 1 (Immediate Convergence): 80–90% of nodes with unambiguous evidence resolve immediately with high confidence (P > 0.99).
  2. Pass 2 (Conditional Diffusion): Dependent nodes receive the locked states of their parents, refining their probability distributions.
  3. Pass 3 (Constraint Satisfaction): Any remaining high-entropy or conflicting nodes are evaluated against formal boundary constraints.

A complex workflow involving 100 interdependent operational decisions resolves in milliseconds, not minutes.


4. Rigorous Calibration and Principled Abstention

If an AI model tells you it is 99% confident, how often is it actually right?

For off-the-shelf generative models, self-reported confidence is notoriously poorly correlated with truth. Models suffer from severe overconfidence on out-of-distribution inputs, happily hallucinating convincing justifications for erroneous outputs.

In industrial, aerospace, and financial applications, an overconfident wrong answer is fatal. A Decision Model must be empirically calibrated:

E[ Indicator(y_pred == y_true) | P(y_pred) = p ] = p

When Kai outputs a probability of 0.85, that decision is empirically correct exactly 85% of the time.

       Uncalibrated LLM (Overconfident)            Calibrated Decision Model (Kai)
  1.0 |                    /                 1.0 |                     /
      |                   /                      |                   /
A 0.8 |             _____/                 A 0.8 |                 /
C     |            /                       C     |               /
C 0.6 |           /                        C 0.6 |             /
      |          /                         |          /
  0.4 |    _____/                          0.4 |        /
      |   /                                      |      /
  0.2 |  /                                   0.2 |    /
      | /                                        |  /
  0.0 +----------------------                0.0 +----------------------
      0.0  0.2  0.4  0.6  0.8  1.0               0.0  0.2  0.4  0.6  0.8  1.0
             CONFIDENCE                                 CONFIDENCE

When entropy exceeds a parameterized risk tolerance τ, the model does not guess. It invokes Principled Abstention:

Entropy(P(y | E)) = - Σ p_i * log(p_i) > τ  ==>  DEFER

The system yields execution back to human specialists or deterministic rulebooks, providing an explicit sensitivity gradient identifying which piece of missing evidence would reduce entropy the most.


5. Replayable Decision Packages (The Audit Trail)

In regulated enterprise environments, "the AI generated this text" is not a valid legal defense or compliance posture.

Every decision evaluated by Kai produces an immutable, cryptographically sealed Decision Package:

{
  "decision_id": "dec_8f0a2c91b4",
  "program_version": "fraud_triage_v2.4",
  "timestamp": "2026-09-30T19:00:00Z",
  "model": {
    "name": "kai-1-core",
    "weights_sha256": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855"
  },
  "evidence_fingerprints": {
    "transaction_context": "sha256:7c9e...",
    "biometric_telemetry": "sha256:4a1d...",
    "historical_risk_profile": "sha256:9f8b..."
  },
  "posterior_distribution": {
    "approve": 0.012,
    "step_up_mfa": 0.976,
    "terminate_session": 0.009,
    "defer_to_fraud_desk": 0.003
  },
  "selected_action": "step_up_mfa",
  "entropy": 0.098,
  "signature": "ed25519:3b7a8f..."
}

If an auditor or incident response engineer queries a transaction years later, they can replay the Decision Package against the exact model revision and verify that the output was mathematically deterministic, tamper-proof, and compliant with policy.


Where Decision Models Belong in Your Stack

Decision Models do not replace generative models. They relieve generative models of the jobs they were never designed to do.

In modern enterprise architectures, systems thrive when tasks are separated cleanly between generation and control:

                          ┌──────────────────────────┐
                          │   Incoming Workload      │
                          └─────────────┬────────────┘
                                        │
                                        ▼
                          ┌──────────────────────────┐
                          │       Enso Router        │
                          │  (Task Intent & Triage)  │
                          └──────┬────────────┬──────┘
                                 │            │
             Generative Tasks    │            │   Control & Decision Tasks
       (Code, Research, Content) │            │   (Routing, Approvals, Safety)
                                 ▼            ▼
                     ┌───────────────┐    ┌───────────────┐
                     │  Zen Models   │    │  Kai Models   │
                     │ (LLM / Reason)│    │ (Decision Mod)│
                     └───────────────┘    └───────┬───────┘
                                                  │
                                                  ▼
                                          ┌───────────────┐
                                          │Typed Decision │
                                          │    Package    │
                                          └───────────────┘
  • Zen: Hanzo's frontier generative reasoning models. Use Zen when you need deep textual research, multi-step chain-of-thought proofs, creative synthesis, and document drafting.
  • Kai: Hanzo's open-weight decision model family. Use Kai when you need deterministic routing, tool selection, safety thresholds, trade execution, triage, and agent loop control.
  • Enso: The high-throughput meta-router that inspects incoming tasks at the microsecond layer, routing generative tasks to Zen and decision tasks to Kai.

Open Weights. Sovereign Infrastructure.

Reliable decision infrastructure cannot be locked inside a closed third-party API that changes behavior without notice. When a model governs production deployments, physical systems, or financial risk, you must have complete sovereignty over its weights, runtime, and data privacy.

Kai 1 is open weights. You can run it:

  1. On the Hanzo Sovereign Cloud: Scaled globally with zero-egress guarantees and sub-millisecond latencies.
  2. In Your Own VPC or On-Premises: Deployable via Docker or Kubernetes on NVIDIA GPUs, AMD ROCm, or high-density CPU nodes.
  3. At the Edge: Fully disconnected for robotics, aerospace, and sovereign defense environments.

Stop asking language models to write essays when your software just needs a decision.

Welcome to the era of Decision Intelligence.


Kai 1 · Open Weights

Decisions, Not Completions

Run Kai through the Hanzo Cloud API today, or download the open weights to deploy calibrated decision models inside your sovereign infrastructure.

Read more