Inference Fleet · atlas

How it's built

Hermes answers through per-host gateways over a Tailscale tailnet. Talos primary is cloud xAI grok-4.6; straylight primary is zai glm-5.3-flash. Fallback and delegation go to straylight llama.cpp :11434; small-model aux (approval, titles, web extract) and Honcho compute go to dixie:11434. Six hosts, two local routers, two clouds of compute.

Hosts 6 Structures 19 Decisions 5 Flows 4
Cloud APIs Connection Gateways Local inference Clients & surfaces Memory plane Other hosts flow
01

Decisions locked

the axes the fleet commits to
AxisDecisionADR
Primary modelCloud is the default; local models never carry the main path. Talos primary is xAI grok-4.6. Straylight primary is zai glm-5.3-flash.
Local routerllama.cpp llama-server on straylight is the main-loop fallback/delegation endpoint. Aux routing (approval/titles/web_extract) is served by dixie:11434.adr/0001-local-router.md
TransportMattermost is the gateway transport and hosts the 'claude' overflow user.
MemoryHoncho runs on rift; its deriver calls dixie:11434 for gen, embed and dialectic.adr/0002-memory-placement.md
GPU ceilingStraylight (Strix Halo) holds models-max 3 under a ~104 GiB TTM cap. Titan holds models-max 2. Dixie holds models-max 3 on 12 GB.adr/0003-gpu-ceiling.md
02

Structures

the 19 things that make up the fleet
03

Representative flows

packet shapes the design implies

Payload shapes are what the design implies, not measured traffic. 4 flows walk the fleet.

04

Context

glossary, one line per noun