lower token variance
Same task, real repo: with Cardumen it costs about 13k tokens each time; without it, one run jumps to 38.7k.
Cardumen does not ask you to trust a perfect demo. The benchmark holds the repo, model and task constant and makes the harness the variable.
Same task, real repo: with Cardumen it costs about 13k tokens each time; without it, one run jumps to 38.7k.
qwen3.5-9b exploring a memory-rich repo: 16 down to 8 exploration steps with Cardumen's code graph.
A public reference that shows scaffolding, not only the model, can change the outcome.
Cross-service audit of the advanced signature flow a real NestJS monorepo · 11 microservices · NATS · frontier model · n=3 per arm. The same complex question, six times: with Cardumen the cost stays in a narrow band; without it, one run jumps.
Precision: It found the real flow bug (idempotency latch) with 6 of 6 file:line citations verified against the code. Precision without hallucination. Cardumen narrows cost and removes the worst case.
The local model runs privately and performs better: it explores less with the code graph, and the harness cuts hallucinations that appear without verified context.
Memory-rich repo. The cost saving shines on local models: they explore inefficiently without reliable context.
In a task with a seeded business rule, a model without Cardumen context rewrote the project convention and declared all tests passed (false). With Cardumen it applied the correct rule and closed against evidence.
Cardumen on/off over the real repo: same model, same prompt, only the harness changes.
The verdict comes from tests, build or an independent adversarial reviewer against the code, not the agent that wrote the answer.
Measure tokens, time, variance and worst case, not only the average or monthly price.
Record where the harness does not help, because that edge defines the right routing.
Bring one real task. We compare the same work with and without Cardumen and return the result, including the limits.