↓ Skip to main content

Abhishek Walia

Streaming architect by day. Toolsmith by compulsion.

Recent

Stop making agents watch each other

Recently I learned about Completion Notifications for agentic work within Codex. I told my Codex agents to stop checking on each other and use completion notifications instead. This stopped the orchestrated set of agents from checking on each other with repeated model turns just to report that the work was still running. This started saving my weekly quota to be used on better things than sleep checks every 10 seconds for a 30 minute pytest run.

I Used One Codex Chat to Orchestrate Another

I just learned that Codex can use one chat to supervise another. These are not subagents, but real threads with one monitoring the other. I am working on a MetaSkill project, and I wanted to try out GPT-6 Astra’s new capabilities to see whether it would work better as an orchestrator or an implementer. The orchestrator experiment is ongoing, and I will tackle the implementer part next. I asked Astra to create a separate chat for the work, assign it the goal, track it, and steer it toward the outcome I wanted. Astra started a Sol chat at Extra High reasoning (exactly as I asked it to), then created a scheduled task in its own thread that checked the worker every ten minutes.

The Design Doc Is the Prompt

AI-assisted development is a multi-pronged thought process. It is a prompting optimization game as well as a model personality/tendency and a big rework problem. The model based on its tendency could decide if it wants to take the prompt too literally or maybe just an outline of what you want and still ignore some pieces that it doesn’t want to care about. GPT 5.5 and Opus 4.7 are prime examples for the divergence in behaviour. Opus is more exploratory in nature while GPT 5.5 takes its instructions pretty seriously (as of today). Once you understand how different models behave, you may unlock the power to work with them as a partner by using the subtle persuasion techniques. I have been dealing with these subtleties of different models for the past 8-10 months of my journey evolving my workflow with them (not the exact same versions, but you catch the drift).

TMUX and agents: Match made in heaven

Claude Code runs in my terminal all day. Although I have started dabbling with Codex, Opencode and others but Claude Code is still the primary thing I’m looking at for hours at a stretch for now. All of those harnesses run on a dedicated VM, not my main machine. Daily backups, isolated environment, nothing else lives there. At some point I realized my terminal config mattered a lot more than it used to.

Sizing a Kafka Cluster from First Principles

Most Kafka sizing advice is hand-wavy. “Start with 3 brokers.” “Scale when you hit a bottleneck.” “Talk to your vendor.” I have spent years on both ends of that conversation, as the engineer asking and the person being asked, and the answer is almost always some version of “it depends, let’s just see how it goes.” That isn’t sizing. Its the easiest way to say “I don’t know.”.