Skip to main content

Tools & resources

This section collects practical implementation support for teams applying the playbook.

Tools are separated from the conceptual methodology pages so readers can first learn the method, then choose the artefact that operationalises it. The groupings below mirror the section's navigation.

WOG products

These are centrally maintained products by GovTech that are available to all agencies as a managed offering.

ItemUse for
LitmusSafety and security testing as a service
SentinelInput and output guardrails

Benchmarks

Fixed datasets that score system behaviour against a common standard, so results can be compared across systems and over time.

ItemUse for
RabakBenchMultilingual safety benchmarking for Singapore context
MinorBenchChild-specific safety benchmarking
Responsible AI BenchmarkComparing application-level safety, robustness, and fairness performance

Guardrails

Runtime classifiers that inspect inputs and outputs and flag or block unsafe content while the system is serving traffic.

ItemUse for
LionGuardLocalised content moderation
Off-Topic guardrailDetecting prompts outside an AI system's intended purpose

Frameworks

Structured methods for deciding what to test and which controls to apply to a system.

ItemUse for
WOG safety testing frameworkStandardised risk taxonomy, metrics, and evaluation protocol for safety testing across agencies
Agentic Risk & Capability FrameworkIdentifying and mitigating risks in agentic AI systems, organised by what the system can do
KnowOrNotGenerating out-of-knowledge-base evaluations to measure hallucination

Tools

Open-source libraries for running evaluations in your own pipeline, and for checking that the judges behind those evaluations are reliable.

ItemUse for
KaleidoscopeAutomated evaluation of AI systems with reliability-scored LLM judges
MetaEvaluatorMeasuring how well LLM judges align with human annotations

Reference

Supporting material for the rest of the playbook: shared definitions, and further reading for teams going deeper.

ItemUse for
External resourcesCurated reading list of papers, benchmarks, and practitioner guides
GlossaryWorking definitions for the terms used across the playbook

Was this page helpful?