The purpose of any test is to get feedback that something is breaking.
And if something breaks, you want to know ASAP.
Once you know it broke, you want to narrow down the exact point where it failed as quickly as possible.
So there are really two time components
- Time to figure out there is an issue.
- Time to find the root cause and fix it.
Let's say you have a pricing matrix based on customer type and order value.
Free + $500 order → 10% discount
Pro + $500 order → 15% discount
Enterprise + $500 order → 20% discount
If someone changes the Enterprise rule from 20% to 15%, a unit test catches that immediately.
You know exactly what failed and where to fix it.
Now imagine catching the same thing through an integration test.
You create an order, call the pricing API, it talks to other services, invoice gets generated, payment flow runs, and at the end you find out the final amount is wrong.
Now you have to figure out whether the issue is in the pricing matrix, API call, invoice calculation, data passed between modules, or somewhere else.
You will eventually find it, but the feedback loop is longer and the search surface is much bigger.
Now the interesting question is - does this change when agents are writing the code?
I don't think the fundamental problem changes.
For an agent also, what matters is -
- How quickly can it know something broke?
- And once it knows, how quickly can it narrow down the issue and fix it?
Yes, agents can debug and fix code much faster than humans.
And if you have time on your side, maybe you don't care. You can let the agent run for hours, execute the entire test suite, debug failures, iterate, and finally release the change.
But I still feel there is a lot of token and compute wastage in doing that if a much faster feedback loop could have caught the same issue early.
So if your integration test runs quickly, gives good logs/traces, clearly points to the failure, and the agent can reliably fix it, then maybe you don't need that unit test.
But if your integration test takes 15 minutes and just says:
"Checkout flow failed"
the agent still has to inspect multiple services, API calls, state changes, logs, etc.
You have still created a much larger search space.
So I don't think there is a black-and-white answer that you should use unit tests or integration tests.
It depends on your system -
- How fast can you run integration tests?
- How good is your observability?
- How quickly can a human or agent narrow down the failure?
- How expensive is it to write and maintain the tests?
- How quickly can the issue be fixed once found?
In my setup, integration tests are still relatively slow.
So I ask my AI to write tests for the request first and then write the code.
Unit tests give the agent an early feedback loop with a narrow failure surface.
Then integration/functional tests verify interaction points and end-to-end behavior.
As integration environments get faster, observability gets better, and agents get better at debugging, maybe we will need fewer unit tests.
So for me, it not unit tests or integration tests?
It is how quickly can your system - human or agent, detect a failure, localize it, and fix it, without wasting unnecessary time, tokens, and compute?