Type something to search...

Engagements

Start small. Land with a corpus. Retain with the refresh.

Every engagement is scoped to your slate and your data. Start with the diagnostic, it credits in full against a corpus engagement, so proving the first number costs you almost nothing.

The door-opener

Diagnostic

1 week

A contamination + harness audit of your existing eval, and a written verdict on what it can and can't tell you.

  • Audit of your current eval setup
  • Contamination + harness risk read
  • Written verdict + next-step scope
  • Credited in full against a corpus engagement
The one that pays for itself Start here

Corpus engagement

end to end

The corpus built, the judge aligned, your slate run, the decision delivered.

  • Private corpus from your production data
  • Judge panel aligned to your experts
  • Full model slate run on your harness
  • Decision table + cost-quality frontier
  • Deployed inside your VPC
The standing answer

Slate refresh

ongoing

The slate churns every few months, and every churn invalidates your last decision. We keep the answer current.

  • Every new model added to the frozen corpus
  • A delta report each cycle
  • Quarterly rubric re-alignment
  • Historical run data compounds on your side

Scoped to your slate and your data. Book a call for a quote, or send a query.

Model cost-down audit

Pay us less than we save you.

We prove whether a cheaper model or a leaner slate holds quality on your actual job, with error bars, so you can switch with evidence instead of guessing. If we can't find savings worth more than the fee, you don't pay it.

How it works →
Book a cost-down audit

Compliance evidence pack

Eval outputs mapped to RBI FREE-AI / EU AI Act / ISO 42001 controls, with an immutable audit trail a risk officer can hand to an auditor.

audit-ready

Enquire →

Independent bake-off

A blind evaluation of shortlisted vendors on your own data, delivered as a procurement scorecard.

vendor selection

Enquire →

Terms The clauses that decide everything

Drafted before the first serious conversation.

You own your corpus

You keep a perpetual, irrevocable licence to the corpus we build from your data. You own your raw source data outright, forever. We retain the framework, tooling, and methodology.

Runs in your VPC

Default posture is deployment inside your own cloud or on-prem. Judge models are called with your keys. Nothing leaves your perimeter.

DPA, drafted first

A data-processing agreement covering DPDP and GDPR, with named sub-processors, retention and deletion terms, and PII minimisation at the boundary.

Liability capped at fees

Cap at fees paid, with carve-outs for data breach and IP infringement, the standard, defensible posture for a services engagement.

Straight answers

The questions engineering leaders actually ask.

Why don't you list prices?

Because every engagement is scoped to your slate, your data, and the job you're deciding, and a headline number would either overstate or understate what you actually need. Book a call and we'll scope it in the first conversation. The diagnostic is deliberately small, and it credits in full against a corpus engagement.

Why not just use promptfoo or DeepEval ourselves?

You can, and you should for quick checks. What those tools don't give you is a corpus built from your production reality, a judge aligned to your experts, clustered statistics, and an independent party whose report a regulator will accept. The tooling is the easy part; the corpus and the trust are what decide anything.

Isn't the token cost of running all these models huge?

No, it's a rounding error. A few thousand items judged by a panel of cheap models costs tens of dollars, not thousands. The real cost is human judgment: building the corpus and aligning the rubric. That's where the value is.

What if the answer is that our current model is already fine?

Then we tell you that, and you've bought certainty cheaply. We measure honestly. An eval that always finds a reason to switch models is selling you something, not measuring anything.

Free · 30 minutes · no commitment

The diagnostic pays for itself.

One week, a written verdict on whether your eval can be trusted, credited in full against the corpus engagement if you continue.