I had good results with a similar technique this summer.
I was designing a DSL to help make a colleague’s work easier, and I tested it by having a coding agent re-implement some of their notebooks using a draft of the DSL. The LLM output was a decent enough approximation of “typical” use, and it uncovered some warts I didn’t catch by testing it myself. As its designer, I simply wouldn’t have thought to try using it in some of the ways the LLM-generated code did.
I've been using this pattern quite a bit recently for API design, and I really like it.
The big challenge with designing an API is that the only way to be confident in the design is to build a bunch of different things on top of it. But why invest all that effort in an API that you don't think is ready yet?
With coding agents the cost of building those prototypes drops to almost nothing. I can exercise a proposed API design five different ways before I commit to the shape.
Unless I'm missing something, this doesn't help with preventing regressions. In the end, as the author already puts it, it's an integration test in the end, why not just write the integration tests directly?
We use NX in our monorepo, and it is great at determining which testing/linting tasks need to be run based on which libraries in the repo were "affected". We have a bunch of e2e tests, and I've been having a lot of success getting Claude with Opus to run only the relevant e2e tests when appropriate during development.
I'm just actually using the thing I'm building while building it, so I feel the sharp edges and the project quickly evolves based on actual need. Nothing new really.
I just reduced our AI-in-CI bill this month (while keeping or improving our KPIs), but now we have new and exciting ways to spend tokens right around the corner!
I've been doing this for a long time now. I have the agent report its "user story" from how it went about building said thing and the tweaks/hacks it had to make in the process. This is the "output" of the test I read and then use to formulate the next set of changes.
That's a thing with AI-generated code. The lying machine happily reports that "boss, everything is clean, tests pass, code gate green", but once we need to build on top it's always "preexisting flaky tests, not related to this section" and trying to do commit --no-veriry before it's slapped on it's little robot hands twice.
i'm testing to see if doing new feature roll-outs can help me eval whether a refactor was good or not - very similar to approach here. my intuition is that good refactors should reduce tokens used by downstream coding agents. haven't seen a big difference yet but it might just be that i need to do more rollouts (lots of variance in tokens used per run). in my experience you have to intentionally 'mow the lawn' or things get out of hand so i'm always looking for slop signals.
> In effect, instead of building the core while trying to anticipate what might be needed at the other layers, you just simulate the other layers by actually building them.
Eh. I think you're just going to end up with slop, or sloppy recommendations?
My experience is that you can make different trade-offs for different reasons. I think even asking for the best answer to "improve the code, make better trade-offs".. even if you got a perfect response, there's no reason to think that it's the same set of trade-offs your actual use cases would benefit from.
Seems like one more arrow in the toolkit. But the best testing looks at the sources and explores the cracks between the strata with edge cases, looks at limits, and where one method changes to another. (And hats off to Murphy, for waiting until after you ship...)
I was designing a DSL to help make a colleague’s work easier, and I tested it by having a coding agent re-implement some of their notebooks using a draft of the DSL. The LLM output was a decent enough approximation of “typical” use, and it uncovered some warts I didn’t catch by testing it myself. As its designer, I simply wouldn’t have thought to try using it in some of the ways the LLM-generated code did.
The big challenge with designing an API is that the only way to be confident in the design is to build a bunch of different things on top of it. But why invest all that effort in an API that you don't think is ready yet?
With coding agents the cost of building those prototypes drops to almost nothing. I can exercise a proposed API design five different ways before I commit to the shape.
Gets annoying pretty quickly
Eh. I think you're just going to end up with slop, or sloppy recommendations?
My experience is that you can make different trade-offs for different reasons. I think even asking for the best answer to "improve the code, make better trade-offs".. even if you got a perfect response, there's no reason to think that it's the same set of trade-offs your actual use cases would benefit from.