← All posts

Why I Prefer Astra to Fable

Long coding-agent runs taught me how to catch drift. Using Astra showed me how much more I could build when I had less of it to catch.

Robot builders assemble an elaborate wooden castle beside a blueprint for a simple box.

The agents had built something I hadn’t asked for, written tests for it, and listed it as a completed objective.

I found out in the final report. Somewhere during the work, the main agent and its sub-agents had talked themselves into design decisions that weren’t in my plan. By the time they told me about it, those decisions had become code, with tests to match. I was reading the summary wondering where any of this had come from.

I’ve run into this with Fable and Opus 5 while building out production systems. They can do good work for a long stretch, then lose the objective, cut a corner, or head off in a direction of their own. Finding that afterwards means sending other agents back through the work to investigate.

I spend all day building this way. My plans routinely run for forty minutes to four hours, usually with at least two going at once. A run that needs untangling takes attention away from everything else I have moving. Enough of those interruptions and I’m spending my time improving the process that was supposed to be doing the building.

Even getting an explanation could be awkward. I’d ask what changed and get dropped into the middle of an argument between agents, complete with numbers justifying why they’d built something one way rather than another. I hadn’t been in that discussion. I’d given them the work and gone off to do something else.

Reminding the agent who it was writing for helped. I’d finally get an explanation I could follow. But even when I warned it beforehand, the same habit kept coming back.

Checking against the plan

What worked for the drift was a separate reviewer after each workload. I gave it the original plan and asked it to inspect the result for whether that plan had actually been done. It didn’t get the builders’ discussion or their summary of what they’d accomplished.

That distinction mattered because the discussion was where the new design had acquired its justification. I wanted the reviewer to check the work against what I’d asked for. Giving it the plan and the result kept that job clear.

This worked a lot better. It also made runs longer. Every workload now had another pass before I could move on, so part of my day was still paying for the agent’s tendency to wander. I’d found an efficient way to use Claude, but it took knowing its habits the way a mechanic knows an engine.

I tend to approach these problems through robotics. That’s my training, and I worked across the stack, from sensors and computer vision to autonomous behaviour. I still see coding agents as robots. They act on a view of the world, and that view can drift away from what’s actually going on.

With a mobile robot, you can estimate where it is by adding up its movement. Small errors accumulate. Better sensors and algorithms help, but over time you need something to correct that estimate, such as a GPS fix or recognising a place you’ve already mapped.

When an agent starts drifting, my instinct is to ask where I can add a GPS. In this case, the original plan gives the reviewer a reference outside the builders’ evolving account of the job. They can agree with each other, produce code, and pass their own tests. I still need to know whether they built what I asked for.

What changed with Astra

Then I tried Astra on basically the same kinds of tasks, and it pushed my projects forward noticeably. It held the larger plan in mind better, followed through on the work, and found lingering problems I hadn’t discovered. I could get a better result without the extra review pass I’d needed with Claude.

When I described Astra as about five times better for me, I meant what that let me accomplish. A few fewer errors let me do a lot more with the rest of my day. I could get further into the build before having to stop, investigate what had been happening, and design another process to manage it. That’s a personal estimate of the gain in my work, with all those interruptions and follow-up jobs included.

I’ve spent much more time with Opus 5 than with Astra, so I know far more about how Claude goes wrong. I still need to go back through more of Astra’s work and find its own quirks. But the improvement in the output was clear enough that I wanted to keep building with it.

That matters particularly to how I work. I think in systems, algorithms, and problems. Agents let me spend more of my time there, choosing an architecture and working through what the software should do. When an agent holds that direction through a long build, I can keep more work moving. When it quietly changes the direction, I have to come back down into the details to find out why.

Keeping what I’ve learned

I’m carrying my existing processes into Astra and refining them as I go. Some checks will be specific to habits I’ve learned from Claude, but the patterns are useful beyond one model. The independent reviewer is one of a toolkit I’ve accumulated over years of using language models for professional work.

I write these processes with the agents as Markdown runbooks and version them in GitHub. That lets me read the instructions, change them, and take them to another agent system. Some come from noticing something wrong in the code. Others come from engineering knowledge I’ve dictated into the process.

The code reads much better now than it did before I put those processes in place. I have agents maintain coding standards, remove duplication, keep files divided by what they do, and review the architecture. Documentation gets attention too, because it’s context the next agent will use to understand the project. It needs to be accurate and helpful.

The runbooks need the same care. Agents helped write them, so they also need periodic auditing and cleanup. I’ve found that a fresh agent, properly briefed on what we’re trying to do, can work through them without much trouble. Being in this all day helps me notice when the instructions need another look.

I’ll use that same approach to learn Astra’s weaknesses. For now, I’m taking the extra room it gives me to build. If your agents keep finishing work that has wandered from the request, try the reviewer I started with: give a fresh agent your original plan and access to the result, then ask where the two differ.

Why I Prefer Astra to Fable
0:00
0:00