case file 01 · open-source · inference platform

hal0

Self-hosted AI inference platform for AMD Strix Halo hardware.

hal0 screenshot
1command install
0cloud dependencies
MITlicence

the problem

Strix Halo machines have the memory to run large models locally, but the software path was a pile of scripts. We wanted a platform you could install, point at a model and trust in production.

what we built

An open-source inference server and control plane: model management, an OpenAI-compatible API, hardware-aware scheduling and a plain web console. Packaged for a single box or a small cluster.

what we'd do differently

Ship the console earlier. The API was solid for months before anyone could see what it was doing.

exhibits

hal0: Agent memory as a web: 120 facts from one bank, linked by time, meaning, cause and entity.
Agent memory as a web: 120 facts from one bank, linked by time, meaning, cause and entity.
hal0: An agent card, flipped to its abilities.
An agent card, flipped to its abilities.
hal0: Inference slots: one model each, with its own port, runtime and live tok/s.
Inference slots: one model each, with its own port, runtime and live tok/s.
hal0: Live telemetry over the unified-memory map: which slot holds which gigabytes.
Live telemetry over the unified-memory map: which slot holds which gigabytes.
hal0: The agent library.
The agent library.
hal0: The tag map of a memory bank.
The tag map of a memory bank.
hal0: The benchmark roster: decode and prefill speed per model, measured on the box.
The benchmark roster: decode and prefill speed per model, measured on the box.
hal0: Stacks: slot, profile and model bundles that load in one step.
Stacks: slot, profile and model bundles that load in one step.
hal0: The model catalog.
The model catalog.
hal0: Detected hardware and backends on a Strix Halo box.
Detected hardware and backends on a Strix Halo box.
hal0: Launch profiles: bench-tuned flags per workload.
Launch profiles: bench-tuned flags per workload.
new project

let's talk.

A 20-minute call. We'll tell you what it would take, roughly what it costs, and whether we're the right shop for it.

01

schedule a 20-minute consult

Tell us two or three times that work for you. We reply within one business day to confirm.

ask for a time
02

send us an email

Goes straight to [email protected].