Josh Johnson
Notes from JJ
Remote Agents, Codex Consumes, and New Model Economics
Why I built my own agent stack, the quiet end of Codex, and Sol / Terra / Luna
This week, I want to talk about why remote agent orchestration is the future, the quiet end of the standalone Codex brand, and why GPT-5.6 Sol/Terra/Luna are changing the economics of building production AI networks.
Remote Agents: Why I Built My Own
The biggest shift in my workflow over the past year hasn't been any single model release — it's been the move from batched prompts to persistent, remote agents that own tasks end-to-end.
I run a custom instance of Hermes Agent as my local base. It's terminal-native, skill-aware, and capable of delegating work across tools and platforms. The real magic happens when you treat each agent as a remote worker with memory, permissions, and a defined role rather than a chatbot you fire questions at.
Why this matters:
- Continuity: agents maintain state across sessions — no more explaining the project from scratch.
- Delegation: one agent can spawn sub-agents for parallel workstreams (code review, research, file ops).
- Integration: agents that can fire Discord messages, run terminal commands, query data warehouses, and draft emails are force multipliers.
If you haven't played with a persistent agent stack, start with Hermes or anything that exposes an MCP/ACP interface. The difference between "chatting with AI" and "managing a workforce" is the difference between a prompt and a process.
Codex Is Dead; Long Live Codex
OpenAI Codex as a standalone branded product is effectively dead. Its capabilities are being folded directly into the ChatGPT brand — voice, screen awareness, and deep code execution will all live inside the main ChatGPT app rather than as a separate CLI experience.
I'm all for it. The fragmentation of "AI coder" tools has confused more people than it's helped. Developers don't need five different terminals to ship work. Integrating strong code execution into the primary ChatGPT product means a single, consistent model personality with less context switching.
The New Model Economics: Sol, Terra, and Luna
We've all been chasing the absolute bleeding edge lately. After Claude Fable and the ultra-heavy next-gen frontier models, I found myself defaulting to the biggest, most expensive brains available for everything.
But a funny thing happens when you actually build and run deep agentic workflows. The reasoning baseline has shifted so high that the workflows are just fundamentally understood.
Enter GPT-5.6 Sol (on medium effort). It's significantly cheaper per token than both Fable and Opus, but it handles the context switching between my local agents, APIs, and data warehouses beautifully.
Terra could probably handle a chunk of my daily pipelines at a fraction of the cost. Luna is where I'd go for the cheapest, fastest reasoning on tasks that don't need full frontier depth.
We aren't just getting "smarter" models anymore; we're getting models that allow us to architect highly cost-efficient, production-grade agent networks.
— Josh
Quick Links
- Hermes Agent: Docs — The open-source agent framework I run locally.
- OpenAI Codex: GitHub — Being folded into ChatGPT.
- GPT-5.6 Family: OpenAI Platform — Compare Sol, Terra, and Luna.