Josh JohnsonNotes from JJ
The Agent Applies. The Model Decides. I Get Time Back.
A Tuesday makeup note about Muse doing real work, JEV making tiny decisions fast, and time with my kids already paying off.

I missed Monday's send, so this is the makeup note. Fortunately, the last week supplied plenty to talk about: one tool turned into a viral meme, another made a Game Boy look like an AI benchmark, and the most important update happened away from the screen.
Muse Found Its Job

Meta Muse has been the AI meme of the week. Invite codes are flying everywhere, and the billion-token offers sound made up even when they are real.
Underneath the referral frenzy, I have found a use case that is genuinely useful: job applications.
I keep a Linear board of roles I am interested in. Muse finds new openings, pulls the descriptions into the board, writes to the requirements, and works through the application. I have been getting through roughly 15 to 20 applications a day—not as a benchmark or a promise, just as the result of giving an always-on agent a real queue instead of another vague prompt.
It still stops for the things that should belong to me: creating a password, answering disability and other EEO questions, or handling something sensitive or ambiguous. Those pauses are not failures. They are the line between useful automation and letting software invent personal answers on my behalf.
Muse may also be the gateway drug for people who never wanted to pay for ChatGPT or learn a pile of AI tools. It feels less like “using AI” and more like handing a task to someone who has a computer.
Yes, referral links are a little cringy. A billion tokens is still a billion tokens, so here is the reveal:
Have fun out there—and remember to touch grass, people. The tools are useful, and the feeling of suddenly being able to do everything is addictive.
JEV Makes the Small Decision Fast

The other model taking over my feed is JEV. I do not have one perfect use case for it yet, and that may be because I first tried to think of it like a tiny LLM. It is not trying to write an essay or hold a conversation. You give it some state and a narrow typed question; it returns a choice, score, or yes-or-no probability that code can use. If you want to experiment, TypeSafe has the technical documentation, and JEV is also available through OpenRouter's Decisions API.
That makes it interesting for the small decisions inside a larger system. I have been testing it in Cerberus, my public-repository code review experiment, to help categorize findings and risk. If you are an AI developer using GitHub, try it on a public repo and tell me where it is useful—or wrong. A classification can organize the review, but it is not proof that code is safe.
The wildest demo I saw was Christian Mathiesen's Pokémon Red experiment. The run finished in 37 hours and 40 minutes across 16,150 decisions. JEV was not staring at pixels and improvising the whole game: a harness read emulator state, presented legal actions and facts, and let the model choose. It wiped 16 times before winning.
That is exactly why the experiment is interesting. The result came from thousands of fast, constrained choices—not one giant chain of reasoning. Games are the fun demo, but the same pattern can fit operations, routing, reviews, and workflows where the options are known and waiting several seconds for a full LLM would be absurd.
JEV does not replace Claude, ChatGPT, or Gemini. It suggests a useful layer between ordinary rules and expensive open-ended reasoning: let the big model build the world, then let a smaller decision model handle the repeated gut checks.
Time Off Is Already Working

The best update has nothing to do with a model release.
I am enjoying the time with my kids. This week my five-year-old was at my desk, vibe coding a My Little Pony and Care Bears mashup sticker game live in Codex. We are learning faster, coding more, and getting closer to the small projects that usually disappear underneath the week. I know more about what they are working on day to day, and they know more about what I am building. The questions travel in both directions now.
That has changed the texture of the work. Coding is not only something I disappear into; sometimes it is something we sit beside each other and do. Their projects pull me out of my usual habits, and mine give them a look at how an idea becomes a system through a lot of imperfect little steps.
There is a funny contrast in spending the week with tools designed to do more without me, then realizing what I most wanted was more room to be present. Productivity only matters if the recovered time goes somewhere worth having.
I do not need a longer experiment to call this part: the time off is already a success.
Quick Links
- Meta Muse: Official announcement — How its dedicated virtual computer, connected apps, and approval controls work.
- JEV documentation: Official introduction — TypeSafe's explanation of choices, scores, probabilities, and confidence.
- JEV on OpenRouter: Model and API page and TypeScript tutorial — Two practical ways to make a first decision call.
- Pokémon Red run: Interactive result and source repository — The harness, measurements, failures, and final run.
- Cerberus: Try a public repository — My browser-based code structure and security review experiment.