Learning and people teams in large organisations have a particular kind of workload. A small number of specialists support thousands of empl…
Buying Microsoft 365 Copilot licences takes an afternoon. Getting a pilot group to trust what Copilot tells them takes months, and most of t…
Benchmark contamination is when benchmark data leaks into a model's training data, so the model has effectively seen the test before taking…
There is no best LLM. There is only the best model for a specific job, measured on your task, through your harness, against your budget. So…
Yes, the harness affects LLM performance, and it affects it a lot. You do not ship a model. You ship a system: a prompt, an agent loop, tool…
LLM-as-a-judge is reliable, but only when you control for its known failure modes. Used as a raw scorer it inherits three documented biases…
Use a public benchmark to survey the field and shortlist models. Use a private eval to decide what you ship. A benchmark is a shared public…
Contamination-free LLM evaluation means testing a model on data it could not have seen during training. If your evaluation set overlaps with…
The pitch your team has heard ten times this year: "We can replace 30% of your ops headcount with AI." The pitch usually comes attached to a…
Most agent demos you'll see are toys. A model wrapped in a loop, given a few tools, and pointed at a sandbox. The interesting question, the…
The first version of any RAG system works on a curated demo set. The user asks a question, the system returns relevant chunks, the model ans…
Most AI products we audit are paying somewhere between 5× and 10× what they need to be paying. The reasons are remarkably consistent across…
Most enterprise AI projects don't fail because the technology doesn't work. They fail because the team picked a workflow where AI doesn't co…
The most common AI quality-assurance setup we see in production is one engineer running a few prompts through the new model version, eyeball…