Josh Johnson
Notes from JJ
The $20 Frontier Strategy, Gemini 3.7 Flash & Grokbot
Why you should max out the "big boys" first, the AI tick-tock cycle, and remote agents
If you read this newsletter, chances are you aren't looking for abstract academic benchmark charts — you're trying to connect with AI on a practical, local, and implementable level.
Over the past couple of weeks, I've talked a lot about model economics, token pricing, and routing workloads to the "right brain for the job." But in talking to builders, engineers, and readers, I've realized something important: premature optimization is confusing when you haven't experienced the raw capability of the frontier.
Before we dive into Google's newly released Gemini 3.7 Flash and xAI's Grok 4.6 / Grokbot, let's break down the single most practical recommendation I can make for anyone looking to get serious with AI.
The $20 Strategy: Max Out the Big Boys First, Optimize Second
If I could give one piece of advice to every builder, analyst, or team leader starting out, it's this: stop pinching pennies on nerfed models before you understand what frontier intelligence can actually do.
Don't stay on free tiers or default to low-compute modes out of habit. Put down the $20/month for a pro license — whether that's ChatGPT Plus, Claude Pro, Gemini Advanced, or X Premium / Grok — and immediately crank the reasoning settings to near or high effort.
Go straight for the heavyweights:
- GPT-5.6 Sol on high reasoning
- Grok 4.6 on high reasoning
- Gemini 3.7 Flash with thinking budget enabled
- Claude Opus 5 and Claude Fable
Push these models hard until you start hitting the platform rate limits. Use that compute to build something real:
- A full-stack web application or internal dashboard
- A working automated email sender and campaign manager
- A local terminal CLI tool that debugs your dev environment
- A multi-step data transformation or ETL pipeline
Why this matters
Until you've watched a top-tier frontier model cleanly solve a 500-line logic bug or orchestrate a complex multi-file refactor on the first attempt, you do not have a realistic expectation of what AI is capable of providing.
Once you've built something tangible with the frontier brains, your baseline expectations permanently reset. You develop real, intuitive pattern recognition for how each model thinks and where their distinct strengths lie.
Only at that point — once you know what top-tier intelligence feels like and start bumping against rate limits — does it make sense to care about efficiency, token costs, and routing smaller tasks down to lighter tiers like Terra, Luna, or standard Flash.
The AI "Tick-Tock" Cycle: Gemini 3.7 Flash Drops
Barely three weeks after rolling out 3.6 Flash, Google officially released Gemini 3.7 Flash on August 13. It introduces native hybrid reasoning with controllable thinking budgets, a 1,000,000 token context window, 65k output limit, and default agentic video processing. Most notably, Google paired this intelligence jump with aggressive introductory pricing ($0.75 / 1M input and $3.75 / 1M output through year-end).
What we're witnessing is the AI frontier settling into a classic hardware Tick-Tock cadence (reminiscent of the legendary Intel and Apple silicon cycles):
- The Frontier Leap (Tick): A lab ships a heavy, boundary-pushing reasoning model (like Opus 5, Sol, or Gemini Pro) that expands the frontier of complex logic and multi-step planning.
- Architectural Compression & Efficiency (Tock): Within weeks, those reasoning capabilities are distilled into hyper-fast, low-cost "Flash" and "Mini" class engines that execute at a fraction of the cost while maintaining high stamina on long-running tasks.
With Gemini 3.7 Flash, the endurance on extended agent loops is the standout feature. When frontier-level reasoning becomes this cheap and reliable over long sessions, you can let background agents run deep repository audits or multi-turn data reconciliations autonomously.
Grok 4.6 & Grokbot: Entering the Remote Agent Arena
Meanwhile, on August 12, xAI rolled out Grok 4.6, following the exact same playbook of expanding capabilities while targeting long-horizon workflows:
- Multimodal Image Generation: New image generation modes with expanded fidelity controls and faster rendering.
- Long-Horizon Agentic Execution: Benchmarked on par with frontier tiers, heavily optimized for multi-step tasks, autonomous code synthesis, and higher abstention (lower hallucination) rates.
- Remote Agent Ecosystem & Grokbot: Open-source persistent frameworks like OpenClaw and community Grokbot harnesses allow users to tether their Grok subscriptions directly into 24/7 autonomous agents across Discord, Telegram, Slack, and terminal daemons.
I haven't put Grokbot through its full paces yet, but testing it out in my homelab is top of the docket for this weekend.
Homelab benchmarks and field notes on Grokbot coming next week.
— Josh
Quick Links
- Gemini 3.7 Flash: Google AI Studio — Hybrid reasoning, thinking budgets, 1M context window, and agentic video execution.
- xAI Grok 4.6: xAI Platform — Long-horizon agentic task execution, codebase management, and new image modes.
- OpenClaw Agent Framework: GitHub — Local-first, persistent personal assistant and multi-platform agent framework.
- Sprout CLI: Sprout Docs — Automated CLI debugging and environment management.