How I stopped being my AI's scheduler
By Arthur O'Keefe, Founder and Chief AI Officer, Bamboo DCM
Published on

I went looking for a better model. What I needed was a better harness.
There is an old military story about a test given to new officers.
The page is almost empty. At the top, a short paragraph describes a flagpole that needs to go up. It gives you the dimensions, the ground, the equipment. Then it asks: you are the junior officer. Write the instructions you would give your sergeant.
All that white space is the trap. It invites you to fill it.
The answer is one line. Sergeant, erect the flagpole.
I heard that story years ago and filed it away as a joke. I have spent the last six months learning that it is the whole thing.
AI made me busier
The first clear sign that our use of AI was working was that I got busier.
That sounds backwards, so let me show you what it looked like. A good prompt shortened a first draft. A repeated task became a skill, and the skill made it consistent. Simple agents let me run a few things at once. Every one of those was a real gain and I felt it.
But I was still the one deciding what happened next. I carried context from one session to another. I checked what came back. I held in my head which thing had to finish before the next thing could start.
My Tuesday looked like this: a few sessions open, two waiting on me, one that had finished overnight that I had not read yet. I moved text from a file into a chat and from the chat back into a file, then lost track of which copy was true. I ate lunch at my desk and called it productivity.
The machines got faster. I got busier. I was still the scheduler, and the scheduler was the bottleneck.

I ran out of tokens
Here is the honest reason I went looking, because it is smaller and more human than the version I first wrote down.
I was burning through my Claude usage. I kept hitting the ceiling. And I wanted a second model in the mix anyway: something cheaper to fall back on, in case the bill or the next release made the decision for me. People around me were talking about other models. New ones kept arriving. I was curious.
There was also a real reason. A lot of my work involves having one system check another one's output, arguing against it and auditing how it went wrong. That only works if the checker is genuinely independent of the thing it is checking. With one model there is nobody independent to ask. Running more than one is the requirement.
So I tried Codex, mostly because other people were using it.
Two days of pain, and then it ran for days
I want to tell this in the order it happened, because the order is the point.
It was a nightmare. I had been working with an enormous context window and had built everything around that assumption. Codex gave me much less room, and almost nothing I had written survived contact with it.
I am glad I pushed through, because it forced me to rebuild. First I had to rethink how my main instruction file worked. Then I had to rethink how it would work with two different models reading it at the same time. Two days of porting and adjusting before both could work on the same material with the same tools. That rebuild deserves its own article and it will get one.
Then it started running.
The first thing that struck me was that it could work on something for a very long time. So I gave it a queue and let it pull.
It ran for days. It sequenced things. It worked out conflicts. It was slower than what I was used to, but I did not have to sit there. That was the moment I understood I had not simply found a tool that could run longer. I had found one that could coordinate.
Then the coordinator became the bottleneck
It was good, and then it was in the way.
One session pulling a queue with a few helpers underneath it is still one lane. I kept catching myself reminding it to do things at the same time instead of one after another. I was writing the flagpole essay again, in a different window: step four before step three, don't forget to check that dependency.
So I stopped handing it pieces of work and started handing it whole projects, each one able to open its own helpers as it went. That is when the shape changed.
It ran both ways, and I should be fair about it. Sometimes it was a loose cannon and I had to build fences. Sometimes it got lazy and I had to work out how to keep it moving. Neither of those is a complaint. Both are what building the bounds actually consists of.
A model and a harness are not the same thing
This is the part I got wrong for a long time, and I think most people have it wrong too.
We talk about models as if choosing one settles how you will work. It does not. The model is the intelligence. The harness is everything around it: how it starts work, what it may do without asking, whether it can run while you sleep, whether a new session can pick up where the last one stopped.
Those are different products, and they are built by different convictions.
Both of these are moving targets, and I am describing them as I found them in 2026, on my own configuration.
The environment I started in is careful. It checks with me. It comes back to ask. For sitting beside someone that is exactly right, and the model itself is superb. But careful is difficult to delegate to, because delegation means the thing keeps going when you are not there.
Codex will do a great deal without asking. That is uncomfortable, and it is the reason it works. Because it will not stop you, you have to decide in advance what it may and may not do. Building those limits is the actual job, not overhead. I spent years around systems where the interlocks were the engineering, not an annoyance bolted on afterwards.
I am describing what I found in my own work, not what anyone intended. But the useful distinction is not which model is smarter. It is how much the harness will let you hand over.
Every time, I get a little less involved
Once something can run inside limits you set, you can start giving it the context you would have used yourself, and it can do the work the way you would have done it.
You do not get there in one move. You get there by being the reviewer instead of the doer, one step at a time, adding a fence wherever you find you needed one.
The useful definition I keep coming back to is that an agent is a model working in a bounded loop on your behalf. It takes a goal and some context, uses the tools it has, looks at what happened, and adjusts until it finishes, hits a stop, or hands control back to you. Everything interesting is in the word bounded. The work is building limits you are comfortable delegating inside, and giving it a way to notice when it is going wrong, not only when it is going well.

The point was never that it answers my email.
The point is that each time round, I am needed a little less. What it learns does not evaporate when the session ends. I step back another pace, the coordinator gets stronger, and the distance between me and the details grows on purpose.
What I do with the time
I want to be careful here, because this is usually where someone tells you the human remains important, and it reads like the human defending its own job.
It does not feel that way at all. It feels like being let go of something.
I still own where this is going and what I am willing to risk. But taking risk is not the same as having an appetite for it. One is something you do; the other is something you have. Getting the coordination off my desk means I can spend my attention on the direction, and on the things I have not tried, instead of on remembering what has to happen before what.
None of this is how we make credit decisions, run compliance, or operate anything a supervisor would call the control environment. That work has its own rules, its own approvals and its own record, and I am not describing it here.
I did not type this article. I talked for a long time, badly and in circles, about everything I disliked about the first draft. The experience is mine, the judgement is mine, the argument is mine, and every word here went past me before it went anywhere. But the organising, the sequencing, the sorting of a pile of dictation into an argument: I directed that and did not do it. The old version of me would have wrestled a stiff draft into shape and settled for it.
Sergeant, erect the flagpole.
Come and ask
I am going to keep writing about this: the rebuild that made two models work side by side, and how you actually give a system enough context to be useful.
If you run a business with one workflow that always seems to need rebuilding from scratch, I would like to compare notes — about the method, not about anyone's book. Come and find me and ask me about the harness.
Arthur O'Keefe is Founder and Chief AI Officer of Bamboo DCM. A computer engineer by training, he has operated a nuclear reactor as a U.S. Navy submarine officer, built financial systems and helped build Movile and iFood as Movile’s Group CFO and Chief Strategy Officer. He builds systems of iteration and writes about the engineering and operating judgment that make them useful. A system of iteration is AI you can push back on. You can run it flat out for a weekend or park it for a week, then find the work waiting where you left it.
About Bamboo DCM
Bamboo DCM is an independent structurer and distributor of corporate and structured credit in Brazil, helping mid-market companies raise capital from institutional investors. It is building toward an agentic credit firm, with systems intended to carry more of the analytical and operating work while domain leaders remain responsible for decisions and verification.