Rendered at 08:35:20 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
julesrms 18 hours ago [-]
It's all very glitzy, but I'm failing to understand how "One prompt in. One result out." is of any importance.
This and other recent AI hype-fests all seem to be obsessed with making agents do more work unattended.
But surely in the real world, anyone who's got a real product to make is going to want to steer what's happening. It's ridiculous to think that anyone with a deadline would write a prompt so perfect that they walk away for 4 days and come back to find the finished product ready to ship.
If you really can write a prompt so complete and perfect that it needs nothing further, then any regular harness could probably also do the job. But if like normal people you need to try something, think about it, iterate, and repeat.. then you also just need a regular harness.
0c3ca83 16 hours ago [-]
> But surely in the real world, anyone who's got a real product to make is going to want to steer what's happening.
The promise of AI is that you won't need to pay people in order to think. A bet that there will be a long term need to steer is also a bet that AI will fail.
The people footing the bills for it are paying because they believe AI will succeed, and there will be no more need for a human in the loop. The reason people are pouring trillions into these AI companies is that they expect we'll make the breakthroughs that make it possible for AI to succeed at automating everything humans are able to do today.
So, skate to where the puck is going, not where it is, and all that.
julesrms 15 hours ago [-]
Until the AIs are running entire companies unattended, then at some point, some human has to transfer some instructions into some actionable representation than the agents can build. That human might one day be a manager rather than a developer, but that transfer of intent is always going to be the problem to solve.
And IMHO while humans are still in charge of these processes, the way that most people can best design and explain what they want is via conversation, exploration, iteration, etc. Not by a single fire-and-forget prompt!
0c3ca83 5 hours ago [-]
Entire companies don't have to be fully unattended; there just has to be enough automation that they're just a small handful of humans only partly in the loop.
Basically, why doesn't the prompt "make me wealthy" work? Find and iterate on the bottlenecks.
J03daSchm0 18 hours ago [-]
The Venn Diagram of people building the AI hype-fests and people building a real product is just two circles
visarga 16 hours ago [-]
> But surely in the real world, anyone who's got a real product to make is going to want to steer what's happening.
In my own harness I log each and every user message by hook and use the model to extract user intent by re-reading the chat log from time to time. The raw messages are very important, they contain information that can be used to refine the harness on the one hand, and to validate if the agent still follows user intent on the other. Models tend to get lost in the details and forget the big picture.
hughw 18 hours ago [-]
Not saying you're wrong. But right now I'm trying to craft some skill prose to instruct an agent how to optimize a certain process based on my own heuristics, and it's failing. Maybe I can let a super agent divine the right skill prose, iterating to see what works. It's worth a try.
julesrms 18 hours ago [-]
OK, but by the time you've worked out the correct prompt to tell your meta-agent how to iterate on optimising the right prompt for your actual agent, perhaps you could have just tried a few things and got it going!
conartist6 18 hours ago [-]
Nice try, but there's a meta-meta-agent that just tries a few things and gets it going while the other agents optimize the prompt.
Or put another way: "You can't fool me, it's turtles all the way down!"
visarga 16 hours ago [-]
You can't fool me - this prompt refinement turtle stack only has 3 turtles. Not much way down.
ssddanbrown 20 hours ago [-]
RSI here is for "recursive self-improvement", instead of a harness being built to help users with repetitive strain injury like I first thought when reading the post title
ginko 1 hours ago [-]
You'd think everyone working in tech would be aware that RSI was a widely used acronym already.
broodbucket 19 hours ago [-]
We gotta do something about these clashing TLAs.
polotics 18 hours ago [-]
Well for starters the use of the term "Recursive" is very dubious.
This is as far as the eye can see all very iterative, there is no tail-call or anything fancy, it is loops. Also "Recursive" I feel kind of tries to imply the LLM's weights are being pushed around in some feedback, because to recurse you have to invoke the thing you're recursing into at its very start right? That would be reinforcement probably, and there is none of that in any such RSI so far, at least not public. Please someone contradict me with examples.
So please: "ISI" for iterative self improvement is fine. Also "ISI" does not fit the (outdated!) Vernon Vinge "singularity" trope, and that is a good thing!
IanCal 18 hours ago [-]
I'm unsure.
The example in the docs of improving nanochat is iterative. It's a looped process in one thing altering a second thing.
What would be recursive is raven updating raven to make it better at doing things. For what I picture as RSI the important part would be that it's able to make itself better at doing things and better at improving itself.
Now there's an "evolver" part that improves the harness over time but I don't know how far that goes or what scope it has to update things.
I guess it's that if you have a function called "optimise" that takes functions and makes them better, calling optimise(my_process) is iterative regardless of how many times you do it. Calling optimise(optimise) is inherently different.
gpm 19 hours ago [-]
Perhaps modelling acronyms in TLA+ will help?
daveguy 18 hours ago [-]
Well, we could add a + after any acronyms after they have a formal definition in TLA+. Except for TLA itself. That would need to be TLA++.
peddling-brink 18 hours ago [-]
I agree, we can’t have all these three letter agencies running amuck.
brookst 19 hours ago [-]
A register is the obvious choice, but of course then you get a gold rush and people squatting on XPZ and stuff without even having invented an opaque term behind it.
flexd 18 hours ago [-]
Time for FLAs?
daveguy 18 hours ago [-]
IIRC there are already FLAs. FLAs too. ROTFL.
wafflemaker 19 hours ago [-]
A very popular yoga kata called sun salutations (available in various versions depending on your fitness/advancement level) helped me get rid of wrist and thumb RSI. It stretches and gently stresses these load bearing ;) joints, thus strengthening them.
Hope this helps the last few folk before searching for RSI will be impossible due to the term being taken over.
For those afraid of Satanism in yoga, there's also non satanic variants where you don't greet each other saying Namaste or say Shanti anywhere during the practice.
brookst 19 hours ago [-]
As long as we’re burying things for archaeologists to find, let it be known that AGI once stood for “adjusted gross income”, a measure used for income taxes.
calebhwin 18 hours ago [-]
Dumb but honest question - do repos like this buy stars? How do they have thousands of stars with very little presence across HN/Reddit/X?
bitpush 17 hours ago [-]
An accusation wrapped as a genuine question. Bravo.
Let me turn this around to you. Do you think your timelines on Reddit/X are indicative of the tech scene of {London,China,India,Indonesia}? What makes you think you know of every popular project out there?
jeffnash 19 hours ago [-]
Reminds me a lot of omnigent (which I am a huge fan of) with a persistent memory layer. Unlike omnigent's subagent threads, the DAG it uses to coordinate other harnesses doesn't look to be durable; I am curious as to whether this is by design or is a forthcoming feature, as this essentially makes or breaks my use case of long-running project-sized implementation sessions.
In any event, it's great to see competition in this meta-harness space, which is likely one that none of the frontier labs will touch since it, by definition, would utilize their competitors' products.
marginalx 18 hours ago [-]
I'm curious if you have a few mins for feedback, what are the top 2 things here that omnigent does that is significantly better for you than latest cc/codex which can launch subagents, auto save memory of a project.
I'm wondering that as these core tools continue to enhance and add these capabilities, how much of a benefit these meta harnesses actually provide.
jeffnash 14 hours ago [-]
I agree the gap is narrowing quickly, especially with workflows in CC. For me, a few large advantages still remain:
1/ Allowing me to easily plug in any harness, using any provider, and make it a first-class worker. Omnigent has out of the box ACP support and it's trivial to use that to add first-class support for any harness out there. I love the ability to have CC + Opus plan, Codex + Luna implement, Pi + Qwen 3.8 give a tie-breaking opinion on a design decision that Opus flagged and Grok and Codex couldn't agree on, all orchestrated by a model of my choosing from any provider using Omnigent's main agent harness.
2/ Reusable agent systems rather than just reusable workflows. You can define agents in YAML whose subagents embody particular roles, with different models/harnesses, skills, plugins, tools, etc. preconfigured for each one.
Of course, claude workflows are now durable but Omnigent's agnt definitions are a bit more abstract in that they define the subagents that are available and how they should work by default rather than the workflow itself (i.e. the specific JTBD). If I have a common workflow that consists of, for example, Sol + Codex writing some script to scrape some data, Pi + a cheap DeepSeek-tier model formatting that data en masse, then Fable + CC doing some advanced analysis on it, I can embody that with a yaml agent definition that I can then use to run with my task of the day as a prompt. All of this is orchestrated by a model of my choice using Omnigent's harness.
This might look like: 'smart scraping agent with all sorts of scraping skills and tools pre-loaded', a 'bulk data processing agent with a cheap, fast model and plenty of pandas/numpy skills preloaded', and 'frontier model to interpret and reason on the implications of the processed data'. The main agent would have instructions about the general workflow of such tasks and when to invoke and delegate tasks to which subagent. The definition describes the workers available to the orchestrator and how they should generally behave, rather than hard-coding the workflow itself. I love that I can create those definitions and re-use them.
All that said, I am sure the labs will come up with their own similar products to (2) (e.g. dots today). I also recently noticed that Claude Code now has subagent 'teams' rather than just 'general-purpose'/'explore' subagents, and these seem to be longer-lived. This seems to be encroaching on the agent yaml definitions, albeit with less fine-grained control on my end. Therefore, the tl;dr (for me at least) is vendor neutrality; I don't think we'll ever see a product coming out of a frontier lab that eagerly delegates a task to their competitor's model (and bank account).
poeticsilence 9 hours ago [-]
if you're interested in this space, is there a good way to contact you? Would love to give you early access to something I think you'll love.
lin7c 18 hours ago [-]
[flagged]
redhale 16 hours ago [-]
I don't want to be too negative, but ... all this for a 0.8% improvement in SWE-bench Verified (90.2 for OpenCode vs 91)? And why is this (saturated) benchmark the one coding benchmark chosen to showcase on the homepage?
Without trying it, this seems like its probably just a massive waste of tokens.
PcChip 14 hours ago [-]
this might be a dumb question, but if it's saturated, then isn't 0.8% improvement a very good improvement?
redhale 13 hours ago [-]
I'm definitely willing to be corrected, but I think it just means that this improvement is close to pure noise -- not good or bad necessarily. Just not a good signal.
To me the really bad sign is that this is the benchmark that they selected to highlight. Why not one of the less-saturated benchmarks where this harness could (in theory) show meaningful improvement over the OpenCode baseline? Seems fishy to me.
arminluschin 20 hours ago [-]
This looks similar to https://paseo.sh/, if I understand correctly. I’ve recently tried it and liked it a lot. Would be nice to see a comparison. When’s the harness of harnesses of harnesses coming?
jfaat 19 hours ago [-]
I'm using paseo heavily too. The main differences here seem to be around how opinionated Raven's orchestration is. They're providing agents, workflows, memory, skills, etc. Paseo gives you some orchestration tools but it's mostly letting the 'native' harnesses do the work. For my workflow, I'm interested in some of what they're doing here but I'm not in buying into their whole system (markdown for memory and calling it RSI, as someone else called out, is not doing it for me). The follow up actions and meta-harness tuning look pretty cool.
They also don't ship an app which is one of the best parts of paseo. Then again Paseo's perf leaves a lot to be desired.
mistercheese 16 hours ago [-]
Yeah I actually don’t even use Paseo’s orchestration much. I’m mostly using Paseo because I want to use Anthropic and OpenAI’s native harnesses for their respective models, but OpenCode for others, with a good remote mobile app experience when I want to check in on agent work or steer things away from my computer.
I haven’t found a better OpenCode remote mobile app experience. If OpenCode or Pi makes one, I might just move to that.
riskable 17 hours ago [-]
I always figured that the free/open weights models like qwen3.8:27b would perform just as well if not better than Claude's latest if you just fed it back into itself enough times. This project seems to prove that this is indeed the case.
What I'd like to see now is how good it can get when you feed the micro models like qwen3.5:0.8b into itself to solve problems. Will it be like toddlers discussing neighborhood politics at a pretend tea party or will it actually get some decent results?
Another game-changer (if this style works out): Just get a model like qwen3.8:27b onto one of those model-on-a-chip cards that makes it 1000x faster and see how fast it can go using the same method.
troyvit 16 hours ago [-]
What you describe reminds me of some studies of jumping spiders. Some aspects of their intelligence (like counting) matches that of a 1 year old human, but because their brains are so tiny it just takes them much longer to do the same counting. IOW it's not the size of the model, but how it's organized and the strategies for using it.
This article has lots of fluff but it describes a lot of what I'm talking about:
Why do none of the projects it built work on my machine? Even the fps game looks cool but can't be downloaded? :0
joshstrange 16 hours ago [-]
I continue to yearn for a harness of harnesses but each one I try (or build) takes me uncomfortably far from the work being done.
I don't want to be a prompt shuttle, though I feel that way sometimes. Performing the same dance for each ticket I work on. My issue is that, to bastardize a common joke/phrase, 50% of the things the agent stops for are things it (or another agent) could answer for me, but it's a different 50% task to task.
With HoH's I constantly feel like I'm getting peppered with unimportant questions or being kept out of the loop of things that really need my eyes on it. Threading that needle has been particularly difficult.
aatd86 20 hours ago [-]
Interesting. Only thing, from someone who has built something similar, is that it moght tend to duplicate certain capabilities that those harnesses handle on their own. Some overlap is bound to happen.
brookst 19 hours ago [-]
This, and also they have to chase changes in model tuning and capabilities. I have a nice little meta-harness focused on product development (requiremwnts, acceptance criteria) for Claude code, and when major new models come it it’s weeks before I can find and fix constraints that are no longer necessary + constraints that have become necessary.
revexos 20 hours ago [-]
Recursion has started. Here we go!
PcChip 20 hours ago [-]
I wish they had also compared it to omp and dsh, I’m curious how it stacks up
monkmartinez 18 hours ago [-]
I think the only way to do that is test yourself... otherwise, you are asking for bias. DSH's "cordis" is very interesting, I haven't tried it yet. I have spent lots of time with pi.dev, omp and hermes. Aspects of all harnesses are great, but there is always something that bugs me.
If you have the chops to evaluate different harnesses, you have the chops to build one that is perfect for you.
mpalmer 20 hours ago [-]
Agents editing their prompts and writing memory to markdown files is not and will never be RSI.
topheroo 15 hours ago [-]
“Built for RSI” sounds like the kind of meaningless thing an AI would throw into a tagline to garner clicks.
fraywing 18 hours ago [-]
It feels like the only thing left people are building are harnesses? Harnesses of harnesses?
yuck39 18 hours ago [-]
Harnesses are only useful when frontier models are incapable of designing their own efficient interfaces with systems which they are improving at rapidly. This will be seen as a transitional artifact of a specific time in the development of general intellegent systems
Kuyawa 19 hours ago [-]
That's one of the most beautiful readmes I've seen in my whole life. Threshold is stunning. I am so impressed even RSI got overshadowed
cyfyifanchen 22 hours ago [-]
We're working on Raven: the harness of harnesses, built for RSI (recursive self-improvement).
Raven brings Claude Code, Codex, and its own Research, Code, Design, and Oncall agents into a shared task graph. The idea is to let different agents handle the parts of a project they are suited to, with shared memory across subagents and context carried across sessions.
The RSI work extends to the harness itself: prompts, policies, strategy code, and playbooks. Raven's specialist harnesses and orchestration layer can be improved independently. Candidate changes are evaluated before adoption. This concerns Raven's own components; it doesn't rewrite Claude Code or Codex internals.
We've used Raven for long-running research and experimentation workflows and for building a Godot game. The repository includes examples and outputs, along with installation instructions. We're also exploring how to develop and refine specialist agents for particular domains.
Raven is pre-alpha and Apache-2.0 licensed. The self-improvement work is experimental; the Curator currently ships in the repository rather than the installed package.
Where do you find coordination between agents breaks down today? We'd also be interested in what evidence you'd want before trusting an agent-generated change to its own harness.
This and other recent AI hype-fests all seem to be obsessed with making agents do more work unattended.
But surely in the real world, anyone who's got a real product to make is going to want to steer what's happening. It's ridiculous to think that anyone with a deadline would write a prompt so perfect that they walk away for 4 days and come back to find the finished product ready to ship.
If you really can write a prompt so complete and perfect that it needs nothing further, then any regular harness could probably also do the job. But if like normal people you need to try something, think about it, iterate, and repeat.. then you also just need a regular harness.
The promise of AI is that you won't need to pay people in order to think. A bet that there will be a long term need to steer is also a bet that AI will fail.
The people footing the bills for it are paying because they believe AI will succeed, and there will be no more need for a human in the loop. The reason people are pouring trillions into these AI companies is that they expect we'll make the breakthroughs that make it possible for AI to succeed at automating everything humans are able to do today.
So, skate to where the puck is going, not where it is, and all that.
And IMHO while humans are still in charge of these processes, the way that most people can best design and explain what they want is via conversation, exploration, iteration, etc. Not by a single fire-and-forget prompt!
Basically, why doesn't the prompt "make me wealthy" work? Find and iterate on the bottlenecks.
In my own harness I log each and every user message by hook and use the model to extract user intent by re-reading the chat log from time to time. The raw messages are very important, they contain information that can be used to refine the harness on the one hand, and to validate if the agent still follows user intent on the other. Models tend to get lost in the details and forget the big picture.
Or put another way: "You can't fool me, it's turtles all the way down!"
This is as far as the eye can see all very iterative, there is no tail-call or anything fancy, it is loops. Also "Recursive" I feel kind of tries to imply the LLM's weights are being pushed around in some feedback, because to recurse you have to invoke the thing you're recursing into at its very start right? That would be reinforcement probably, and there is none of that in any such RSI so far, at least not public. Please someone contradict me with examples.
So please: "ISI" for iterative self improvement is fine. Also "ISI" does not fit the (outdated!) Vernon Vinge "singularity" trope, and that is a good thing!
The example in the docs of improving nanochat is iterative. It's a looped process in one thing altering a second thing.
What would be recursive is raven updating raven to make it better at doing things. For what I picture as RSI the important part would be that it's able to make itself better at doing things and better at improving itself.
Now there's an "evolver" part that improves the harness over time but I don't know how far that goes or what scope it has to update things.
I guess it's that if you have a function called "optimise" that takes functions and makes them better, calling optimise(my_process) is iterative regardless of how many times you do it. Calling optimise(optimise) is inherently different.
Hope this helps the last few folk before searching for RSI will be impossible due to the term being taken over.
For those afraid of Satanism in yoga, there's also non satanic variants where you don't greet each other saying Namaste or say Shanti anywhere during the practice.
Let me turn this around to you. Do you think your timelines on Reddit/X are indicative of the tech scene of {London,China,India,Indonesia}? What makes you think you know of every popular project out there?
In any event, it's great to see competition in this meta-harness space, which is likely one that none of the frontier labs will touch since it, by definition, would utilize their competitors' products.
I'm wondering that as these core tools continue to enhance and add these capabilities, how much of a benefit these meta harnesses actually provide.
1/ Allowing me to easily plug in any harness, using any provider, and make it a first-class worker. Omnigent has out of the box ACP support and it's trivial to use that to add first-class support for any harness out there. I love the ability to have CC + Opus plan, Codex + Luna implement, Pi + Qwen 3.8 give a tie-breaking opinion on a design decision that Opus flagged and Grok and Codex couldn't agree on, all orchestrated by a model of my choosing from any provider using Omnigent's main agent harness.
2/ Reusable agent systems rather than just reusable workflows. You can define agents in YAML whose subagents embody particular roles, with different models/harnesses, skills, plugins, tools, etc. preconfigured for each one.
Of course, claude workflows are now durable but Omnigent's agnt definitions are a bit more abstract in that they define the subagents that are available and how they should work by default rather than the workflow itself (i.e. the specific JTBD). If I have a common workflow that consists of, for example, Sol + Codex writing some script to scrape some data, Pi + a cheap DeepSeek-tier model formatting that data en masse, then Fable + CC doing some advanced analysis on it, I can embody that with a yaml agent definition that I can then use to run with my task of the day as a prompt. All of this is orchestrated by a model of my choice using Omnigent's harness.
This might look like: 'smart scraping agent with all sorts of scraping skills and tools pre-loaded', a 'bulk data processing agent with a cheap, fast model and plenty of pandas/numpy skills preloaded', and 'frontier model to interpret and reason on the implications of the processed data'. The main agent would have instructions about the general workflow of such tasks and when to invoke and delegate tasks to which subagent. The definition describes the workers available to the orchestrator and how they should generally behave, rather than hard-coding the workflow itself. I love that I can create those definitions and re-use them.
All that said, I am sure the labs will come up with their own similar products to (2) (e.g. dots today). I also recently noticed that Claude Code now has subagent 'teams' rather than just 'general-purpose'/'explore' subagents, and these seem to be longer-lived. This seems to be encroaching on the agent yaml definitions, albeit with less fine-grained control on my end. Therefore, the tl;dr (for me at least) is vendor neutrality; I don't think we'll ever see a product coming out of a frontier lab that eagerly delegates a task to their competitor's model (and bank account).
Without trying it, this seems like its probably just a massive waste of tokens.
To me the really bad sign is that this is the benchmark that they selected to highlight. Why not one of the less-saturated benchmarks where this harness could (in theory) show meaningful improvement over the OpenCode baseline? Seems fishy to me.
They also don't ship an app which is one of the best parts of paseo. Then again Paseo's perf leaves a lot to be desired.
I haven’t found a better OpenCode remote mobile app experience. If OpenCode or Pi makes one, I might just move to that.
What I'd like to see now is how good it can get when you feed the micro models like qwen3.5:0.8b into itself to solve problems. Will it be like toddlers discussing neighborhood politics at a pretend tea party or will it actually get some decent results?
Another game-changer (if this style works out): Just get a model like qwen3.8:27b onto one of those model-on-a-chip cards that makes it 1000x faster and see how fast it can go using the same method.
This article has lots of fluff but it describes a lot of what I'm talking about:
https://knowablemagazine.org/content/article/mind/2021/are-s...
I don't want to be a prompt shuttle, though I feel that way sometimes. Performing the same dance for each ticket I work on. My issue is that, to bastardize a common joke/phrase, 50% of the things the agent stops for are things it (or another agent) could answer for me, but it's a different 50% task to task.
With HoH's I constantly feel like I'm getting peppered with unimportant questions or being kept out of the loop of things that really need my eyes on it. Threading that needle has been particularly difficult.
If you have the chops to evaluate different harnesses, you have the chops to build one that is perfect for you.
Raven brings Claude Code, Codex, and its own Research, Code, Design, and Oncall agents into a shared task graph. The idea is to let different agents handle the parts of a project they are suited to, with shared memory across subagents and context carried across sessions.
The RSI work extends to the harness itself: prompts, policies, strategy code, and playbooks. Raven's specialist harnesses and orchestration layer can be improved independently. Candidate changes are evaluated before adoption. This concerns Raven's own components; it doesn't rewrite Claude Code or Codex internals.
We've used Raven for long-running research and experimentation workflows and for building a Godot game. The repository includes examples and outputs, along with installation instructions. We're also exploring how to develop and refine specialist agents for particular domains.
Raven is pre-alpha and Apache-2.0 licensed. The self-improvement work is experimental; the Curator currently ships in the repository rather than the installed package.
Code and examples: https://github.com/EverMind-AI/Raven
Where do you find coordination between agents breaks down today? We'd also be interested in what evidence you'd want before trusting an agent-generated change to its own harness.
Of course, it's impossible to know for sure what was LLM processed or not, but this post got classified that way.