This isn't even a problem: the author is just confused why Anthropic didn't bake all of Claude Code documentation into the harness, but of course it makes no sense to pollute context like that.
> Agents suck at communicating plans
True, communication could be better.
> Agents take any excuse to stop working
Just use /loop.
> Agents are only useful when they take unnecessary risks
Just use a sandbox. OK, technically, this one isn't literally solved by Anthropic/OpenAI, but it is solved by a million agent sandbox startups.
Just using subagents was a pretty useful gain. My only issue is sometimes I forgot to tell the agent to use them. Probably need a hook or skill to do it so I don't need to remember.
I think you mean "sandbox" like what is in the Codex TUI. Parent poster meant a VM or container (or bubblewrap/firejail) that doesn't let the agent to edit files outside of specific paths or run dangerous commands, so you can turn the whole confirmation off.
what other sandbox do you need besides running a harness in a container with the right folder mounted and some due diligence like -nonewprivileges? are we talking host network sandboxes here? I think he's just referring to a rogue agent running rm -fdr --no-preserve-root, and that's safe in a container.
Ask your favorite LLM what a malicious script dropped into the right place in your .git directory can do to you the next time you run git outside of the sandbox.
And of course, if the agent has access to the network (which is probably required in order to talk to the AI server), then you have to think about what other things on your local network it could get into.
I keep running into this. It's nice seeing others here struggle. I guess when all your doing is one-shot simple trivial tasks who cares. I've found they're great for that. But once I need to do real work, everybody makes their own esoteric abandonware that kinda works but kinda doesn't. Stars aren't a perfect indicator, but I haven't seen anything over a few hundred for these things I'm finding on GitHub. Same with downloads of plugins. It was very isolating feeling like I'm the only one not enamored by the state of this.
Same here. I have a feeling that once the AI hype cycle blows over to a degree (not that it will go away, just calm down), we'll be left with an enhanced version of pre-boom development. LLMs are hitting a dead end. I'm not convinced they can be bandaid-fixed to become AGI or SI or whatever the cool new trendy term is. They're good at many things, but sustaining a lean and directed software project's architecture over long periods of time doesn't seem to be one of them. That still remains to be further proven over time, but I've already seen plenty of examples of this.
Also, "esoteric abandonware" is a great way to put it. I keep looking for solutions to problems, finding dead projects that haven't been committed to in months, they have all the indicators of AI slopification (the primary one for me being emoji-filled READMEs). I don't think these types of projects will be around for the long-haul. These projects are just as much about community and people as they are about code. And people are much less likely to care about a short-lived project that doesn't have the backing of people that intuitively understand the internals of it. It's like building on sand.
Personally, I've decided that allowing the "agents" (these buzzwords are quickly becoming pet peeves of mine) to write the code for me is a bad idea, because the faster the LLM constructs architecture than I can, the less I'm absorbing the details of it, and the less likely I am to intuit obvious shortcomings. They do help me work much faster than before though, because I still use them like a research/learning assistant chatbot. Learning what I'm doing wrong faster, while still being the one putting the pieces together, is very great.
I see your point, but even the fact that I can speak to my computer and anything useful happens is still a miracle to me.I don't think I will ever get accustomed to how good the new models are.I just can't keep up.And I do this for a living.Ten hours a day.
A lot can be improved, but this is already so much speed.
I used to relate to this article quite a bit. In the last 3-4 months, not so much. I've found that the latest models -- Opus 5.5, Astra, etc. juggle multiple tasks, delegate exceptionally well and are very good at working independently.
I still occasionally have issues with open-weight models, but the frontier labs have solved the above for most use cases.
the agent never stops and says, “Wait, this is something another model could do cheaper and faster.” It just plows on with the slow, expensive model. Conversely, the agent never says, “This model is too dumb for this task. Let me tag in a smarter one.”
This is a pretty big failing, which is compounded by the fact that most humans don't know which model to pick, or make assumptions based on Anthropic's hierarchy or "effort" involved.
Like Fable: your toughest challenges. You mean, like Fields Medal toughest challenges? Or analyzing and updating three monster spreadsheet toughest challenges? Or writing a new novel in the style of William Gibson toughest challenges?
OTOH, would you trust the vendor to pick the best model for you? Would you trust them not to prioritize their own load-balancing concerns first?
The descriptions are near-useless and tend to flip around, as model families are not released in sync anymore, that's true, but fortunately, thanks in a big way to subscription pricing, the choice is simple: start with the best model on offer, and when you run out of quota, downgrade to the next best (or briefly switch providers).
I have some significant experience in context engineering, but I'm most familiar with Codex as a coding harness right now.
Sam Altman said "you don't need to write prompts anymore". This has widely been panned as something someone selling AI would say. If you care about the output and how much time/money it costs to produce it.. Just giving Codex an abstract tasks with zero extra guidance isn't going to produce the best results..
Codex Astra can do a great-(ish) job as a project coordinator dispatching tasks to a pool of 6.1 Sol sub agents. You can even give it an explicit goal and ownership over ensuring the work is carried out efficiently.
However the OOTB harness(and prompt) configuration may not do this for you. You'll have to provide guidance over how you want it to operate through your prompt, a skill, or etc.
And I'll say even though it's really good at this.. Having even more layers than 2 can help; a single agent given too many responsibilities will start to become fixated on a number of them while neglected others. You can check in occasionally to "nudge" it or you might need to split out responsibilities more..
I will say it's crazy Codex doesn't have more built-in task and sub agent management features. I almost wish that it had some stock orchestration patterns that worked OOTB, and then you could opt-in to a leaner setup where you provide more of the instruction.
I double checked the publish date before writing down this comment. Because I use the same harness (Opencode, to be specific) as the author, and some of the features are just right there. Like opencode can start multiple subagents for different tasks in parallel. Also I usually ask the main model (e.g. opus 5.5) to pick subagent models for me, and it has no problem identify the difficulty of work and delegate large portion of them to gpt-luna.
I didn't want to complain about OpenCode too much because it's open-source and kind of the scrappy underdog to Claude/Codex, but I do find its multitasking support pretty limited. I have to specifically remind it to spin up subagents, and when it does, it'll do 2-3 tasks in parallel and then wait until all of them finish. It can't seem to inspect tasks in progress, and if a subagent dies or hangs, the main agent can't seem to retrieve any information from it, even though I as the human can look at the session and see what happened.
Are you doing anything special to make OpenCode multitask well?
> Is there a better coding agent for me?
I’ve only tried Claude, Codex, OpenCode, Cline, and Pi.
I started to gloss over hard after a few paragraphs because I don't really have any the problems you describe anymor; at least not to the point i'd be blogging about them. Instead, I've just been iterating with an agent on various pi extensions that solve the issues.
As with operating systems - you can hold out in the hope somebody solves the right mix of issues in general, and they match your situation well-enough.
That's your choice.
I would note though that being the passive consumer gets you either Mac or Windows UX & prices.
Came here to write exactly that. Especially about Pi, which they claim they've used.
Isn't this the point of Pi to implement all the features you need yourself? Complaining about lack of features in Pi doesn't make sense to me, it's their whole identity.
Although I must say that describing specific problems is valuable on its own. It's just (a) why stop there and (b) if stop there, why shape it as a complaint and not as a list of features that are worth discussing.
I'm confused how you see this article as a complaint about Pi. I only mention Pi at the very end to say it's one of the agents I've tried. I checked it out, found it too barebones for me and moved on, but I understand why some people prefer it. The existence of Pi isn't a counterargument to, "Why don't agents support development workflows that should be commonplace?"
It always seems bizarre to me that people complain about software now. You can just write the thing you want. If you want to do things in parallel do them in parallel. Claude Code will allow you to run multiple instances in the same folder and let them communicate. Previously I used to let them intermediate through a communication bus but now they seem to be able to talk to each other.
I let most agents work asynchronously and don't pay attention so I don't care that much about the sequential nature. But if it's a problem for you then fix your harness. This is a bit like saying "Why are shoes so shit? There's a stone in one and it just gets stuck there and your foot steps on it and it hurts". Take off the shoe, and shake out the rock. Put the shoe back on. You have the power.
I've now been trying for six months to get agents to produce decent code. As in readable / easy to follow, easy to change. I'm doing something really wrong. It's killing me. Everything it produces will work, buts it diabolically over complicated. I've built skills that have helped. But not massively. I've used other people's skills, in particular Matt pococks grill me and Dex hortlys show me. These have helped a bit. I work in enterprise, I want to be proud of what I'm producing, but trying to understand and then cajole and agents code into something that's good is exhausting. If anyone here has been through this and can share how they got through this, I'd really appreciate the help.
Edit: I have access to codex, vscode, GitHub co pilot cli and all anthropic and openai models (excluding mythos).
Same, it’s so smart but so dumb. I explain an architecture change to it and write examples of strongly typed Go code and how to store structure, it agrees and then proceeds to write some untyped string map ball of mud that has 900 specializations.
But maybe that’s on us, AI doesn’t care about all these special cases, it’s not debt to it as it will simply read them
all when making changes. We’re obsessed with quality and what code is supposed to look like but those are human standards, AIs evolve to look at this complexity as a single picture, they can simply see through it so what is spaghetti code to us is merely some code to them that works as it should and is efficient. It’s interesting we can see how the two things drift apart, you would think at some point AI generated code should explode but it hold together unreasonably well in most cases…
Use hooks to run a bunch of review steps after -all- code writing steps your agent does and give it the exact review criteria you just described (via git hooks, or your agent harness of choice's own hooks eg https://code.claude.com/docs/en/hooks or AGENTS.md). So after every step where the agent writes the shitty untyped string map ball of mud, your orchestrator/main thread agent that spawned the code-writing-subagent spawns a follow-up review agent automatically that is given that output, your prompt that explains what well written Go code looks like, and even the sample/golden-path code of your choice to use as a style guide.
Each time you encounter a shitty thing you hate, add a new 'review type' / 'thing to watch out for' and just ask your agent to add it to your hooks for you. This works well with Claude at least.
I have about a dozen or so hooks that run on every integration branch my agents write that review for all sorts of things from correctness to spec, performance improvement opportunities, modularity, analysis of any dependencies added, 'definition of done', UI/UX, etc.
I recently told Claude it should run the whole suite of reviews twice. I will probably go on and proceed to having it run like 5 times eventually idfk.
But the more you start asking your agents to modify their own behavior, using the native solutions offered by Cursor, or Claude, or Codex, the sooner you'll start to feel better about the results.
I'm having bad results with SotA models. Some questions, if you're up for it:
In your workflow, who implements the review feedback - the review subagent or the code-writing-subagent?
Do have a baseline styleguide (like Google's Go style guide) for the review subagents, or is it entirely the subjective things and specific corrections? I remember 6 months ago it seemed like piling general "good taste" code advice into AGENTS.md was considered bad.
Do you move between harnesses or have you gone all in on claude? I've bounced between claude/codex/omp, maybe to my detriment.
Yeah I regularly have 5+ agents working in 1 repo simultaneously even modifying the same file.
Since they were all running on my singular local machine.
Builds and tests were interfering at 1 point so 1 automatically proposed writing a single script that basically mutex locked it with appropriate wait and timeouts.
It had even added this to my own personal agentic todo backlog. They've proposed new skills for me.
They even have split my decisions to human decisions. Proposed and approved work.
They can iterate on approved work without me just fine.
And I only just started with agentic coding in last few weeks before that I was mostly a copy paste chat person.
But you didn't just write the thing you wanted, you're relying on other software that happens to do it, and if you were happy with the communication bus you would've stuck with that.
Your shoe analogy also breaks down because really the shoe is the issue, not the stone. And expecting everybody to make their own shoes is, well, I mean we just don't do it that way anymore for good reason. Let the cobblers make the shoes, and the runners wear them.
AI models can multitask/use parallel subagents just fine; the issue is with harnesses that don't make it a priority via the default system prompt.
I run an LLM server with Qwen 3.6 in the office, and OpenCode, which the OP mentioned, usually defaults to sequential TODO lists, and it works fine with our little LLM server with 3-4 parallel users. But I noticed that once in a while the LLM got overloaded with requests in the queue, and you couldn't do anything for 20-30 minutes. My investigation led me to an employee who used QwenCode. I tried it myself then, and indeed, it immediately launched something like 6 parallel subagents, where OpenCode would have sequential TODOs with the same model by default.
So in the end, I had to detect QwenCode on the server side and serialize all its parallel requests into a single request queue, because it made life miserable for other OpenCode users :)
One thing that seems to help for me is to do the docs before plans (collaboratively edit with agent). Then I understand what this change is going to look like from the user's perspective before we start implementation. This seems to help keep things on track.
While I don't use this plugin a lot anymore, I think doc-driven development is one of the most effective ways to do development in any paradigm, I should probably refresh this plugin and use it more:
The reason Claude looks up the docs of itself is because the model doesnt know about harness features that havent landed yet. The harness itself can be ahead of the model. Happens all the time.
I think definitionally the model can't know about the harness, since the model is at least a few months out of date. So every model must be behind it's harness or something really weird is going on.
Makes me wonder about doing some archaeology and trying out really old harnesses on modern models...
You know "who" could implement all those feature in a new harness? Claude! So why don't you ask it to do it, open-source it, and let all of us bask in the brilliance of you multi-tasking agent?
Actually, the last time I asked Claude Code about itself, it located and read its own minified source and told me something that wasn’t even in the docs.
It's true! It feels like we've been talking about harness optimisation and 'cool features' available in the cli tools for months at this point, but ostensibly there has not really been any significant upgrades to these harnesses since at least Claude Code imo. It does feel like a contrived way to harvest more and more information and test each conversation/action tool, to the detriment of those of us actually using them!
This is a great list for future Agent / Harness software engineers to read!
Oh sure, there may be some Agents/Harnesses that already accomplish some of these things -- but there doesn't seem to be one (as of the present day that I write this) that accomplish all of them...
As someone that watches the Agent/AI Harness (and related software) space, I will definitely be referring back to, and re-reading this list in the future!
I asked DeepSeek to translate a page to five languages and it opened five subagents each one working independently on the translation, once they all finished the main agent informed me of the job completion with a bell. Fantastic!
All these problems are solved.
> Agents can’t manage tasks
Say "use subagents".
> Agents can’t delegate
Say "use subagents that are Haiku/Sonnet"
> Agents have never heard of agents
This isn't even a problem: the author is just confused why Anthropic didn't bake all of Claude Code documentation into the harness, but of course it makes no sense to pollute context like that.
> Agents suck at communicating plans
True, communication could be better.
> Agents take any excuse to stop working
Just use /loop.
> Agents are only useful when they take unnecessary risks
Just use a sandbox. OK, technically, this one isn't literally solved by Anthropic/OpenAI, but it is solved by a million agent sandbox startups.
Just using subagents was a pretty useful gain. My only issue is sometimes I forgot to tell the agent to use them. Probably need a hook or skill to do it so I don't need to remember.
Funny the sandbox is a real security benefit but it just turned into an extra confirmation for me. I'd rather have it than not.
I think you mean "sandbox" like what is in the Codex TUI. Parent poster meant a VM or container (or bubblewrap/firejail) that doesn't let the agent to edit files outside of specific paths or run dangerous commands, so you can turn the whole confirmation off.
what other sandbox do you need besides running a harness in a container with the right folder mounted and some due diligence like -nonewprivileges? are we talking host network sandboxes here? I think he's just referring to a rogue agent running rm -fdr --no-preserve-root, and that's safe in a container.
Ask your favorite LLM what a malicious script dropped into the right place in your .git directory can do to you the next time you run git outside of the sandbox.
And of course, if the agent has access to the network (which is probably required in order to talk to the AI server), then you have to think about what other things on your local network it could get into.
I keep running into this. It's nice seeing others here struggle. I guess when all your doing is one-shot simple trivial tasks who cares. I've found they're great for that. But once I need to do real work, everybody makes their own esoteric abandonware that kinda works but kinda doesn't. Stars aren't a perfect indicator, but I haven't seen anything over a few hundred for these things I'm finding on GitHub. Same with downloads of plugins. It was very isolating feeling like I'm the only one not enamored by the state of this.
Same here. I have a feeling that once the AI hype cycle blows over to a degree (not that it will go away, just calm down), we'll be left with an enhanced version of pre-boom development. LLMs are hitting a dead end. I'm not convinced they can be bandaid-fixed to become AGI or SI or whatever the cool new trendy term is. They're good at many things, but sustaining a lean and directed software project's architecture over long periods of time doesn't seem to be one of them. That still remains to be further proven over time, but I've already seen plenty of examples of this.
Also, "esoteric abandonware" is a great way to put it. I keep looking for solutions to problems, finding dead projects that haven't been committed to in months, they have all the indicators of AI slopification (the primary one for me being emoji-filled READMEs). I don't think these types of projects will be around for the long-haul. These projects are just as much about community and people as they are about code. And people are much less likely to care about a short-lived project that doesn't have the backing of people that intuitively understand the internals of it. It's like building on sand.
Personally, I've decided that allowing the "agents" (these buzzwords are quickly becoming pet peeves of mine) to write the code for me is a bad idea, because the faster the LLM constructs architecture than I can, the less I'm absorbing the details of it, and the less likely I am to intuit obvious shortcomings. They do help me work much faster than before though, because I still use them like a research/learning assistant chatbot. Learning what I'm doing wrong faster, while still being the one putting the pieces together, is very great.
A perfect polished artifact is not important. Every person can just vibe their own good enough solution. And that’s fine.
We live in a society though
I see your point, but even the fact that I can speak to my computer and anything useful happens is still a miracle to me.I don't think I will ever get accustomed to how good the new models are.I just can't keep up.And I do this for a living.Ten hours a day.
A lot can be improved, but this is already so much speed.
I used to relate to this article quite a bit. In the last 3-4 months, not so much. I've found that the latest models -- Opus 5.5, Astra, etc. juggle multiple tasks, delegate exceptionally well and are very good at working independently.
I still occasionally have issues with open-weight models, but the frontier labs have solved the above for most use cases.
the agent never stops and says, “Wait, this is something another model could do cheaper and faster.” It just plows on with the slow, expensive model. Conversely, the agent never says, “This model is too dumb for this task. Let me tag in a smarter one.”
This is a pretty big failing, which is compounded by the fact that most humans don't know which model to pick, or make assumptions based on Anthropic's hierarchy or "effort" involved.
Like Fable: your toughest challenges. You mean, like Fields Medal toughest challenges? Or analyzing and updating three monster spreadsheet toughest challenges? Or writing a new novel in the style of William Gibson toughest challenges?
No, but it will do the opposite You can choose Sonnet and set "/advisor opus" and it will reach up when it thinks it needs to.
OTOH, would you trust the vendor to pick the best model for you? Would you trust them not to prioritize their own load-balancing concerns first?
The descriptions are near-useless and tend to flip around, as model families are not released in sync anymore, that's true, but fortunately, thanks in a big way to subscription pricing, the choice is simple: start with the best model on offer, and when you run out of quota, downgrade to the next best (or briefly switch providers).
I have some significant experience in context engineering, but I'm most familiar with Codex as a coding harness right now. Sam Altman said "you don't need to write prompts anymore". This has widely been panned as something someone selling AI would say. If you care about the output and how much time/money it costs to produce it.. Just giving Codex an abstract tasks with zero extra guidance isn't going to produce the best results..
Codex Astra can do a great-(ish) job as a project coordinator dispatching tasks to a pool of 6.1 Sol sub agents. You can even give it an explicit goal and ownership over ensuring the work is carried out efficiently.
However the OOTB harness(and prompt) configuration may not do this for you. You'll have to provide guidance over how you want it to operate through your prompt, a skill, or etc.
And I'll say even though it's really good at this.. Having even more layers than 2 can help; a single agent given too many responsibilities will start to become fixated on a number of them while neglected others. You can check in occasionally to "nudge" it or you might need to split out responsibilities more..
I will say it's crazy Codex doesn't have more built-in task and sub agent management features. I almost wish that it had some stock orchestration patterns that worked OOTB, and then you could opt-in to a leaner setup where you provide more of the instruction.
I double checked the publish date before writing down this comment. Because I use the same harness (Opencode, to be specific) as the author, and some of the features are just right there. Like opencode can start multiple subagents for different tasks in parallel. Also I usually ask the main model (e.g. opus 5.5) to pick subagent models for me, and it has no problem identify the difficulty of work and delegate large portion of them to gpt-luna.
OP here.
I didn't want to complain about OpenCode too much because it's open-source and kind of the scrappy underdog to Claude/Codex, but I do find its multitasking support pretty limited. I have to specifically remind it to spin up subagents, and when it does, it'll do 2-3 tasks in parallel and then wait until all of them finish. It can't seem to inspect tasks in progress, and if a subagent dies or hangs, the main agent can't seem to retrieve any information from it, even though I as the human can look at the session and see what happened.
Are you doing anything special to make OpenCode multitask well?
I believe OpenChamber [0] ticks all or at least almost all of TFA's boxes.
[0] https://openchamber.dev/
OP here.
Looks like this is a layer on top of OpenCode. Seems interesting. I'll check it out. Thanks for the tip!
It does use OpenCode under the hood to run the agents but I think it's quite a bit more than a layer on top of it.
The dev is very responsive on Discord and I'm sure he'd be happy to hear your thoughts and suggestions!
> Is there a better coding agent for me? I’ve only tried Claude, Codex, OpenCode, Cline, and Pi.
I started to gloss over hard after a few paragraphs because I don't really have any the problems you describe anymor; at least not to the point i'd be blogging about them. Instead, I've just been iterating with an agent on various pi extensions that solve the issues.
As with operating systems - you can hold out in the hope somebody solves the right mix of issues in general, and they match your situation well-enough.
That's your choice.
I would note though that being the passive consumer gets you either Mac or Windows UX & prices.
Came here to write exactly that. Especially about Pi, which they claim they've used.
Isn't this the point of Pi to implement all the features you need yourself? Complaining about lack of features in Pi doesn't make sense to me, it's their whole identity.
Although I must say that describing specific problems is valuable on its own. It's just (a) why stop there and (b) if stop there, why shape it as a complaint and not as a list of features that are worth discussing.
OP here.
I'm confused how you see this article as a complaint about Pi. I only mention Pi at the very end to say it's one of the agents I've tried. I checked it out, found it too barebones for me and moved on, but I understand why some people prefer it. The existence of Pi isn't a counterargument to, "Why don't agents support development workflows that should be commonplace?"
Skill issue, sorry.
It always seems bizarre to me that people complain about software now. You can just write the thing you want. If you want to do things in parallel do them in parallel. Claude Code will allow you to run multiple instances in the same folder and let them communicate. Previously I used to let them intermediate through a communication bus but now they seem to be able to talk to each other.
I let most agents work asynchronously and don't pay attention so I don't care that much about the sequential nature. But if it's a problem for you then fix your harness. This is a bit like saying "Why are shoes so shit? There's a stone in one and it just gets stuck there and your foot steps on it and it hurts". Take off the shoe, and shake out the rock. Put the shoe back on. You have the power.
I've now been trying for six months to get agents to produce decent code. As in readable / easy to follow, easy to change. I'm doing something really wrong. It's killing me. Everything it produces will work, buts it diabolically over complicated. I've built skills that have helped. But not massively. I've used other people's skills, in particular Matt pococks grill me and Dex hortlys show me. These have helped a bit. I work in enterprise, I want to be proud of what I'm producing, but trying to understand and then cajole and agents code into something that's good is exhausting. If anyone here has been through this and can share how they got through this, I'd really appreciate the help.
Edit: I have access to codex, vscode, GitHub co pilot cli and all anthropic and openai models (excluding mythos).
Same, it’s so smart but so dumb. I explain an architecture change to it and write examples of strongly typed Go code and how to store structure, it agrees and then proceeds to write some untyped string map ball of mud that has 900 specializations.
But maybe that’s on us, AI doesn’t care about all these special cases, it’s not debt to it as it will simply read them all when making changes. We’re obsessed with quality and what code is supposed to look like but those are human standards, AIs evolve to look at this complexity as a single picture, they can simply see through it so what is spaghetti code to us is merely some code to them that works as it should and is efficient. It’s interesting we can see how the two things drift apart, you would think at some point AI generated code should explode but it hold together unreasonably well in most cases…
Use hooks to run a bunch of review steps after -all- code writing steps your agent does and give it the exact review criteria you just described (via git hooks, or your agent harness of choice's own hooks eg https://code.claude.com/docs/en/hooks or AGENTS.md). So after every step where the agent writes the shitty untyped string map ball of mud, your orchestrator/main thread agent that spawned the code-writing-subagent spawns a follow-up review agent automatically that is given that output, your prompt that explains what well written Go code looks like, and even the sample/golden-path code of your choice to use as a style guide.
Each time you encounter a shitty thing you hate, add a new 'review type' / 'thing to watch out for' and just ask your agent to add it to your hooks for you. This works well with Claude at least.
I have about a dozen or so hooks that run on every integration branch my agents write that review for all sorts of things from correctness to spec, performance improvement opportunities, modularity, analysis of any dependencies added, 'definition of done', UI/UX, etc.
I recently told Claude it should run the whole suite of reviews twice. I will probably go on and proceed to having it run like 5 times eventually idfk.
But the more you start asking your agents to modify their own behavior, using the native solutions offered by Cursor, or Claude, or Codex, the sooner you'll start to feel better about the results.
I'm having bad results with SotA models. Some questions, if you're up for it:
In your workflow, who implements the review feedback - the review subagent or the code-writing-subagent?
Do have a baseline styleguide (like Google's Go style guide) for the review subagents, or is it entirely the subjective things and specific corrections? I remember 6 months ago it seemed like piling general "good taste" code advice into AGENTS.md was considered bad.
Do you move between harnesses or have you gone all in on claude? I've bounced between claude/codex/omp, maybe to my detriment.
Yeah I regularly have 5+ agents working in 1 repo simultaneously even modifying the same file. Since they were all running on my singular local machine. Builds and tests were interfering at 1 point so 1 automatically proposed writing a single script that basically mutex locked it with appropriate wait and timeouts. It had even added this to my own personal agentic todo backlog. They've proposed new skills for me.
They even have split my decisions to human decisions. Proposed and approved work. They can iterate on approved work without me just fine.
And I only just started with agentic coding in last few weeks before that I was mostly a copy paste chat person.
But you didn't just write the thing you wanted, you're relying on other software that happens to do it, and if you were happy with the communication bus you would've stuck with that.
Your shoe analogy also breaks down because really the shoe is the issue, not the stone. And expecting everybody to make their own shoes is, well, I mean we just don't do it that way anymore for good reason. Let the cobblers make the shoes, and the runners wear them.
AI models can multitask/use parallel subagents just fine; the issue is with harnesses that don't make it a priority via the default system prompt.
I run an LLM server with Qwen 3.6 in the office, and OpenCode, which the OP mentioned, usually defaults to sequential TODO lists, and it works fine with our little LLM server with 3-4 parallel users. But I noticed that once in a while the LLM got overloaded with requests in the queue, and you couldn't do anything for 20-30 minutes. My investigation led me to an employee who used QwenCode. I tried it myself then, and indeed, it immediately launched something like 6 parallel subagents, where OpenCode would have sequential TODOs with the same model by default.
So in the end, I had to detect QwenCode on the server side and serialize all its parallel requests into a single request queue, because it made life miserable for other OpenCode users :)
Enjoyed this article, lots I can relate to.
One thing that seems to help for me is to do the docs before plans (collaboratively edit with agent). Then I understand what this change is going to look like from the user's perspective before we start implementation. This seems to help keep things on track.
While I don't use this plugin a lot anymore, I think doc-driven development is one of the most effective ways to do development in any paradigm, I should probably refresh this plugin and use it more:
https://github.com/tmpdir-org/tmpdir-claude-code-marketplace...
The reason Claude looks up the docs of itself is because the model doesnt know about harness features that havent landed yet. The harness itself can be ahead of the model. Happens all the time.
I think definitionally the model can't know about the harness, since the model is at least a few months out of date. So every model must be behind it's harness or something really weird is going on.
Makes me wonder about doing some archaeology and trying out really old harnesses on modern models...
They are as dumb as their instructions.
Have you tried telling models about your dream agent environment?
They can build it.
You know "who" could implement all those feature in a new harness? Claude! So why don't you ask it to do it, open-source it, and let all of us bask in the brilliance of you multi-tasking agent?
Actually, the last time I asked Claude Code about itself, it located and read its own minified source and told me something that wasn’t even in the docs.
So that's how model distillation is done. Just ask for the source code directly.
agents have progressed a LOT in the last year. Claiming they're still bad in the same way they were bad in 2025 is inaccurate.
It's true! It feels like we've been talking about harness optimisation and 'cool features' available in the cli tools for months at this point, but ostensibly there has not really been any significant upgrades to these harnesses since at least Claude Code imo. It does feel like a contrived way to harvest more and more information and test each conversation/action tool, to the detriment of those of us actually using them!
>https://mtlynch.io/why-are-coding-agents-so-dumb/#what-i-wis...
This is a great list for future Agent / Harness software engineers to read!
Oh sure, there may be some Agents/Harnesses that already accomplish some of these things -- but there doesn't seem to be one (as of the present day that I write this) that accomplish all of them...
As someone that watches the Agent/AI Harness (and related software) space, I will definitely be referring back to, and re-reading this list in the future!
An excellent post!
Perhaps is not the agent that is dumb?
I asked DeepSeek to translate a page to five languages and it opened five subagents each one working independently on the translation, once they all finished the main agent informed me of the job completion with a bell. Fantastic!
Sooo, which agent?
> My dream agent
I have no idea what stops that person from just making it, with an agent of course.
> If I ask Claude how to use the features of Claude, it has to search online to figure out what this “Claude” thing is.
As expected. A model's knowledge is what was it ingested a creation.t
Unfortunately what we get is worse - for the same reason. Model version thinks it is its previous version.