Tools & resources
This section collects practical implementation support for teams applying the playbook.
Tools are separated from the conceptual methodology pages so readers can first learn the method, then choose the artefact that operationalises it. The groupings below mirror the section's navigation.
WOG products
These are centrally maintained products by GovTech that are available to all agencies as a managed offering.
Benchmarks
Fixed datasets that score system behaviour against a common standard, so results can be compared across systems and over time.
| Item | Use for |
|---|---|
| RabakBench | Multilingual safety benchmarking for Singapore context |
| MinorBench | Child-specific safety benchmarking |
| Responsible AI Benchmark | Comparing application-level safety, robustness, and fairness performance |
Guardrails
Runtime classifiers that inspect inputs and outputs and flag or block unsafe content while the system is serving traffic.
| Item | Use for |
|---|---|
| LionGuard | Localised content moderation |
| Off-Topic guardrail | Detecting prompts outside an AI system's intended purpose |
Frameworks
Structured methods for deciding what to test and which controls to apply to a system.
| Item | Use for |
|---|---|
| WOG safety testing framework | Standardised risk taxonomy, metrics, and evaluation protocol for safety testing across agencies |
| Agentic Risk & Capability Framework | Identifying and mitigating risks in agentic AI systems, organised by what the system can do |
| KnowOrNot | Generating out-of-knowledge-base evaluations to measure hallucination |
Tools
Open-source libraries for running evaluations in your own pipeline, and for checking that the judges behind those evaluations are reliable.
| Item | Use for |
|---|---|
| Kaleidoscope | Automated evaluation of AI systems with reliability-scored LLM judges |
| MetaEvaluator | Measuring how well LLM judges align with human annotations |
Reference
Supporting material for the rest of the playbook: shared definitions, and further reading for teams going deeper.
| Item | Use for |
|---|---|
| External resources | Curated reading list of papers, benchmarks, and practitioner guides |
| Glossary | Working definitions for the terms used across the playbook |
Was this page helpful?