AI & Unit Testing Verbosity
08-22-2026
At work, I spent a day or so updating a npm package that contained a bunch of peer dependencies. I have been using AI more deliberately and less often as a means of upskilling myself rather than training LLMs.
When I did reach for Cursor, I noticed a behavior that compelled me to type up a quick post, mostly so I remember to investigate my hypothesis.
What I found was that agents consistently run entire frontend + backend testing suites to “validate” very small increments of work, which obviously wastes a lot of time, and even worse, I suspect that unit test output is bloating context windows and leading to a polluted context for the agent, as well as a waste of money (tokens… tokens are money!). Even with prompting that I think was adequate enough, but maybe not a spec-driven directive, I wouldn’t expect Java unit tests to run based on npm changes.
Even with a bunch of specific rules.mdc files that define rules to:
- etc
- only run the entire testing suite during a final verification step
- only run unit tests for affected files/modules
- only test certain behaviors based on product requirements
I have a few ideas about this, some of which I think are rational and some are honestly paranoid thoughts that assume the worst of the technocrat overlords that run frontier AI shops:
- Companies have every motivation ($) to overemphasize code output metrics, and unit tests are probably the area most likely to be accepted into the codebase. After all, they’re helpful, prevent bugs, and are unlikely to affect production, right? (Not right)
- More code = more context pulled into the context window = higher token usage.
- Ineffective prompting: unit tests can really be deprioritized when there are other factors at play, like timelines and requirements that frequently change. Prompting write unit tests for these new files/functions/methods generally produces LOTS of slop.
This is all just half-brained speculation based on my experience recently. I’m going to investigate this more by doing some experiments and reading more articles and research papers. I’m not a gambler, but would bet that this becomes a bigger problem as codebases mature with AI-driven code development.
Further Reading
I’ve read some of these, but just collecting some resources for when I come back to this.
- Agentic test processes, LLM benchmarks, and other notes on agentic coding from Galapagos Island
- AI-Assisted Unit Test Writing and Test-Driven Code Refactoring: A Case Study
- An Empirical Evaluation of Using Large Language Models for Automated Unit Test Generation
- An Empirical Study of Unit Test Generation with Large Language Models
- A Review of Large Language Models for Automated Test Case Generation
- Benchmarking LLMs for Unit Test Generation from Real-World Functions
- Budget-aware Test-Sufficient Context Selection for Java Unit Test Generation
- Do Coverage and Mutation Scores of LLM-Generated Test Suites Correlate with Their Effectiveness?
- Do influence tactics matter? investigating prompt framing effects in LLM code generation
- Don’t Let an LLM Write Your Unit Tests
- Effective test generation using pre-trained Large Language Models and mutation testing
- Impact Assessment of Structured Results for the Reliability of LLM-generated Tests
- Prompt engineering in LLMs for automated unit test generation: A large-scale study (some interesting links to Java-based unit test resources in this one)
- Stop Vibe Coding Your Unit Tests
- Test case generation using large language models: a systematic literature review