If you run AI in a Singapore financial institution, you validate LLMs and agents by producing evidence that maps to the MAS FEAT principles and IMDA’s governance guidance. FEAT sets the expectations. Independent, contamination-free evaluation with error bars and an audit trail supplies the proof those expectations ask for.

TL;DR

  • Singapore’s MAS issued the FEAT principles (Fairness, Ethics, Accountability, Transparency) for AI and data analytics in finance. These are principles, not a hard AI Act, so you show adherence through evidence rather than clause-by-clause conformity.
  • The Veritas initiative provides an assessment methodology and toolkits to help firms apply FEAT in practice.
  • IMDA published a Model AI Governance Framework, with added guidance for Generative AI, that complements FEAT at the national level.
  • A MAS-regulated firm needs evidence that is credible to an outside reviewer: independent testing, contamination-free corpora, statistically honest results, and a reproducible audit trail.
  • One rigorous evaluation, run in your own VPC, can produce artifacts that answer FEAT and map to global frameworks at the same time.

What is MAS FEAT and where does evaluation fit?

FEAT is a set of principles for the responsible use of AI and data analytics in Singapore’s financial sector, organized around Fairness, Ethics, Accountability, and Transparency. It is guidance oriented, so it tells you the qualities your AI use should have rather than handing you a checklist to tick.

Evaluation is where those qualities become demonstrable. Fairness and ethics are claims until you test the system on representative tasks and can show how it behaves. Accountability means someone can point to who validated the model and how. Transparency means the basis for a deployment decision is recorded and reviewable. Each principle rests on the same thing underneath: honest measurement you can show to someone outside the build team.

That is the gap a lot of internal testing leaves open. A model can pass a demo and still lack the evidence FEAT expects, because that evidence is about credibility and reproducibility, not a passing score.

How does the Veritas initiative help apply FEAT?

Veritas is the industry initiative that turns FEAT from principles into practice by providing an assessment methodology and toolkits.

The value for an AI governance lead is that Veritas gives you a structured way to ask the FEAT questions of a specific system, rather than reasoning about fairness or transparency in the abstract. It helps you frame what to assess and how to document the assessment.

What it does not do for you is generate the underlying evidence. A methodology tells you which questions to answer. You still need a rigorous evaluation to produce defensible answers, and for LLMs and agents that means testing on tasks that reflect real use, on data the model has not seen, with results you can trust. Veritas frames the assessment. The evaluation fills it.

What does IMDA’s GenAI governance guidance add?

IMDA’s Model AI Governance Framework sets national-level guidance for responsible AI, and its added guidance for Generative AI addresses the risks specific to models that generate open-ended output.

For a financial institution, the two operate together. FEAT and Veritas are the finance-sector view; IMDA’s framework is the broader national posture that GenAI-specific guidance extends. If you deploy an LLM or an agent, you are working within both, and both point at the same underlying need: evidence that the system was tested seriously and governed transparently.

Generative systems raise the bar because their failure modes are harder to see. A model that answers fluently can still be wrong, and an agent that touches money can take the wrong action while appearing to work. That is exactly why the evaluation has to be independent and contamination-free, not a self-graded internal run.

What evidence does a MAS-regulated firm actually need?

You need evidence that would convince a reviewer who did not build the model. That reframes the whole exercise, because most internal testing is built to convince the team that already believes.

Concretely, that means a defined scope drawn from real workflows, a held-out corpus the model has not trained on, results reported with error bars rather than a single average, a defensible scoring method, a record of who ran the test and when, and a deployment decision with confidence attached. Independence matters because a builder grading their own build cannot supply the accountability FEAT asks for. A contamination-free corpus matters because a score on data the model has already seen measures memorization, not capability.

We describe how we assemble that evidence in our method. The short version is that the evidence core is not framework-specific. It is just rigorous evaluation, and FEAT draws from it.

How does one eval map to FEAT and global frameworks?

The same evaluation artifacts answer FEAT principles and the questions other regimes ask, so you build the evidence once and map it. Treat the framework descriptions here as directional, and confirm the exact expectations with your own compliance and legal teams.

Eval artifact (produced once)FEAT principleIMDA GenAI guidanceGlobal frameworks
Independent evaluator, separate from builderAccountability, EthicsResponsible governance of GenAIIndependence of validation
Contamination-free, held-out corpus on real tasksFairness in measured behaviorTesting suited to generative riskAccuracy and robustness testing
Results with error bars, not point estimatesTransparency of true performanceHonest reporting of model limitsDefensible validation methodology
Trajectory and tool-selection scoring for agentsAccountability for actions takenOversight of autonomous behaviorValidation of actual system behavior
Documented scope, rubric, and scoring methodTransparency of the basis for claimsDocumentation of governanceTechnical and validation documentation
Timestamped, reproducible audit trailAccountability you can review laterTraceability of decisionsRecord-keeping and supervisory review
In-VPC execution, data never leaves your controlFairness and data governanceData control for GenAI systemsData residency and governance

One row, several columns. You produce the left-hand artifact once, and it answers a FEAT question, an IMDA question, and a global-framework question at the same time. We walk through the cross-framework version of this in one eval, four frameworks.

Why run the evaluation inside your own VPC?

Because data control is part of the evidence, not a separate concern. FEAT’s fairness and governance expectations, and IMDA’s guidance for generative systems, both care about how the data underlying your models is handled.

Running the evaluation in your own VPC means model weights, prompts, and customer data never leave your control. You get the independent, contamination-free result without shipping sensitive material to a third party, which satisfies data governance expectations at the same time as the performance ones. We cover the mechanics in in-VPC LLM evaluation.

FAQ

Is FEAT a regulation we have to pass like an audit? FEAT is a set of principles, not a hard AI Act with clauses to certify against. You demonstrate adherence through evidence and governance rather than a pass or fail conformity test. That is why credible, independent evaluation matters so much: it is how a principle becomes something you can show.

Can our internal team produce this evidence themselves? You can generate strong internal signal, but you cannot generate independence by definition, and accountability under FEAT leans on a claim that is separate from the builder. An external, contamination-free evaluation supplies the credibility a self-scored run cannot.

Does mapping to FEAT mean we ignore SR 11-7 or the EU AI Act? No. The evidence core is shared, so a rigorous eval that satisfies FEAT also produces most of what those regimes ask for. Start with what applies to you and reuse. See one eval, four frameworks.

What is different about evaluating agents rather than a single model? Agents take actions, so you evaluate the trajectory and tool selection, not just the final answer. For money-touching flows that distinction is the whole point, and it is what accountability under FEAT depends on.

You do not need a new evaluation for every principle or every framework. You need one rigorous, independent, contamination-free eval whose artifacts map to FEAT, IMDA guidance, and the global regimes you also answer to. To see which of your obligations one evaluation could cover, book a free eval diagnostic and we will map it to your specific LLMs and agents.