Josh Johnson

Notes from JJ

Week 09 August 1, 2026 · 7 min read

Multi-Harness Architecture & Strategic Model Economics

Inside Nora OS, 27-minute meeting optimization, and AI Basics 101

This week, I'm pulling back the curtain on my operational multi-agent stack at murderszn.github.io/labs. We're breaking down why the endgame of AI adoption is ruthless optimization — and closing with a zero-fluff AI crash course covering the 6 core terms every builder and leader needs to know.

Inside Nora OS: Multi-Harness Operational Architecture

Nora OS multi-harness operational architecture

Most people setting out to use AI get stuck treating models like a glorified Google search box — typing prompts into a single chat window and manually copying text back and forth.

To run real operations, automate complex workflows, or build production systems, you have to move past single-chat windows to a multi-harness operational architecture.

I recently published My Team & Tools — murderszn × Nora OS to map out exactly how my AI operational stack is structured:

  1. The Operator (Vision & Strategy): That's me. I set system architecture, strategic priorities, and final execution oversight.
  2. Primary Work Interface (Claude Code): Client delivery, warehouse transformations, complex engineering. Connected into Databricks, Jira, and GitHub CI/CD.
  3. Chief of Staff & Always-On Ops (Hermes Agent / Kimi 2.7): Runs 24/7 on Gmail triage, Discord daemons, ComfyUI, homelab telemetry, and nightly GitHub backups.
  4. Multi-Harness Router: Routes across Pi Harness, Antigravity CLI, Grok CLI, and Codex CLI.

When you give each AI brain a defined role, a dedicated harness, and a clear set of tools, you stop chatting with AI and start operating a force-multiplying workforce.

Strategic AI Economics: Luna, Terra, and the Optimization Mindset

Model pricing and routing architecture

If there is one operational truth I keep coming back to: the endgame of AI adoption is ruthless optimization.

Over the past year, we've seen dramatic drops in model pricing alongside massive capability gains in smaller, faster models like Luna and Terra. Yet I still see enterprise users throwing frontier models at every mundane prompt.

Let's be honest: throwing heavy frontier models at simple tasks is a massive waste of tokens, compute energy, usage limits, money, and time.

Frontier models are incredible when you need deep architectural reasoning, complex math, or creative heavy lifting. But let's face reality — the meeting is in 27 minutes. You don't have 90 seconds to wait while a massive frontier model overthinks a standard data cleanup.

  • Frontier models: Complex code architecture, deep logic, high-stakes system design.
  • Terra & Luna: Route 80% of daily execution — email triage, doc summaries, API formatting, pair programming — to the smallest model that reliably works.

Stop overpaying for compute you don't need. Optimize your model routing for speed, context efficiency, and cost.

AI Basics 101: A Crash Course in 6 Essential Terms

AI vocabulary and key terms reference card

If you've been reading about AI for months but still feel on the fence — let's strip away the buzzwords.

  1. Model (The Brain): The underlying neural network trained to recognize patterns, reason, and generate text, code, or media.
  2. Token (The Currency): The fundamental unit of text an LLM processes (~4 characters or ~0.75 words). Costs, speed, and context limits are all measured in tokens.
  3. Context Window (Working Memory): Input context is what you load in. Output limit is the maximum response in one turn.
  4. Harness (The Steering Wheel): The wrapper CLI or app that connects you to the raw LLM — prompt engineering, environment access, execution controls.
  5. Skill (The SOPs): Standardized instructions loaded so the model performs a task the same way every time.
  6. Agent (The Hands): A model paired with real tools that can plan multi-step actions and execute independently.

— Josh