I have been going through a bit of a rough patch, or my entries would be more frequent.
Long-term / Long-context
One of the questions this project attempts to describe, if not answer, is what an agent might do when given more agency and a longer time frame. The major use case for which LLMs and generative AI is being marketed is single-session, single-feature work; I wanted to find out what kinds of capabilities, pitfalls, traps, etc. might be inherent in a longer-term objective.
“Emergent Behavior”
This phrase is fraught in the AI community. I won’t dive into that; suffice to day, I use it here with tongue firmly in-cheek. That said, the behaviors exhibited by an agent given a longer context window are interesting, and obviously not apparent when using them for shorter, simpler tasks. That makes sense; presumably (presumably!) the teams who have developed them have determined (and/or optimized towards) the use case(s) for which they are being advertised.
The techniques required for an agent to be useful over long periods of time form more of an art than a science at this point. There is no accepted procedure; there is barely a best practice. It’s the wild west. A context window is a context window, and when it’s full, it’s full. Some mechanic(s) must be in place to allow the agent to reason over material much larger in scope than 1,000,000 tokens.
It’s the behavior resulting from those mechanics which is interesting. Or at least, has been interesting to me, so far. More on that later. One technique I’ve been considering to help along these lines is employing multiple agents.
Two is (always?) better than one
One technique often described to handle managing longer contexts (like that of a project) for agents to handle (whose scope is considerably smaller) is that of the “handoff.” An agent summarizes the work it’s completed that session, and the next agent, with a freshly empty context window, reads it to get up to speed and (purportedly) continue where the previous one left off.
That’s as it may be; there are better and worse ways to do it, and helpful and not-so-helpful techniques that go alongside it. But what if we add some parties to the equation? Instead of the user and a series of agents, what happens if we create a relay of agents? A team of agents? A hierarchy of agents?
That’s what I’ve been considering of late. Agents handing off sessions to each other, or one level up in the chain. There are a lot of challenges before you even sit down to write a line of code, and not all of them are technical. Cloud-based LLMs are expensive. User agreements can be vague.
But that’s what I’m up to at the moment. More as it develops.