Josh Johnson
Notes from JJ
Multi-Harness Architecture & Strategic Model Economics
Inside Nora OS, 27-minute meeting optimization, and AI Basics 101
This week, I'm pulling back the curtain on my operational multi-agent stack at murderszn.github.io/labs. We're breaking down why the endgame of AI adoption is ruthless optimization — and closing with a zero-fluff AI crash course covering the 6 core terms every builder and leader needs to know.
Inside Nora OS: Multi-Harness Operational Architecture
Most people setting out to use AI get stuck treating models like a glorified Google search box — typing prompts into a single chat window and manually copying text back and forth.
To run real operations, automate complex workflows, or build production systems, you have to move past single-chat windows to a multi-harness operational architecture.
I recently published My Team & Tools — murderszn × Nora OS to map out exactly how my AI operational stack is structured:
- The Operator (Vision & Strategy): That's me. I set system architecture, strategic priorities, and final execution oversight.
- Primary Work Interface (Claude Code): Client delivery, warehouse transformations, complex engineering. Connected into Databricks, Jira, and GitHub CI/CD.
- Chief of Staff & Always-On Ops (Hermes Agent / Kimi 2.7): Runs 24/7 on Gmail triage, Discord daemons, ComfyUI, homelab telemetry, and nightly GitHub backups.
- Multi-Harness Router: Routes across Pi Harness, Antigravity CLI, Grok CLI, and Codex CLI.
When you give each AI brain a defined role, a dedicated harness, and a clear set of tools, you stop chatting with AI and start operating a force-multiplying workforce.
Strategic AI Economics: Luna, Terra, and the Optimization Mindset
If there is one operational truth I keep coming back to: the endgame of AI adoption is ruthless optimization.
Over the past year, we've seen dramatic drops in model pricing alongside massive capability gains in smaller, faster models like Luna and Terra. Yet I still see enterprise users throwing frontier models at every mundane prompt.
Let's be honest: throwing heavy frontier models at simple tasks is a massive waste of tokens, compute energy, usage limits, money, and time.
Frontier models are incredible when you need deep architectural reasoning, complex math, or creative heavy lifting. But let's face reality — the meeting is in 27 minutes. You don't have 90 seconds to wait while a massive frontier model overthinks a standard data cleanup.
- Frontier models: Complex code architecture, deep logic, high-stakes system design.
- Terra & Luna: Route 80% of daily execution — email triage, doc summaries, API formatting, pair programming — to the smallest model that reliably works.
Stop overpaying for compute you don't need. Optimize your model routing for speed, context efficiency, and cost.
AI Basics 101: A Crash Course in 6 Essential Terms
If you've been reading about AI for months but still feel on the fence — let's strip away the buzzwords.
- Model (The Brain): The underlying neural network trained to recognize patterns, reason, and generate text, code, or media.
- Token (The Currency): The fundamental unit of text an LLM processes (~4 characters or ~0.75 words). Costs, speed, and context limits are all measured in tokens.
- Context Window (Working Memory): Input context is what you load in. Output limit is the maximum response in one turn.
- Harness (The Steering Wheel): The wrapper CLI or app that connects you to the raw LLM — prompt engineering, environment access, execution controls.
- Skill (The SOPs): Standardized instructions loaded so the model performs a task the same way every time.
- Agent (The Hands): A model paired with real tools that can plan multi-step actions and execute independently.
— Josh
Quick Links
- My Team & Tools: murderszn.github.io/labs — Interactive Nora OS hierarchy and harness map.
- Sprout CLI: Sprout Docs — Automated environment setup and CLI debugging.
- Sprout Discord: Join the community — Share feedback and agent workflows.