Running 1 ALE - Trust Layer evaluation š 1 Explore and verify AI benchmark scores with a Trust Layer
GUIDE: Governed Unified Intelligence for Document-to-Artifact Generation in Enterprise Settings Paper ⢠2608.12133 ⢠Published Aug 12
SAGE: Governed Artifact Generation from Enterprise Guidelines Paper ⢠2609.17775 ⢠Published 26 days ago
RL-ADA: A World-Feedback Framework for Adversarially Robust Enterprise Dialogue Agents Paper ⢠2609.02902 ⢠Published Jul 6 ⢠1
GVD: Governed Versioning and Deduplication for Document Repositories Paper ⢠2609.17696 ⢠Published 26 days ago
Running Agents Personal Assistant Benchmark š§ Scores a personal assistant by what it did on the device
Voice or Stereotype? Disentangling Acoustic and Content-Based Gender in Speech-to-Speech Models Paper ⢠2609.09263 ⢠Published Sep 8
SEABED: SouthEast Asian Benchmark for Evaluating Audio Reasoning Paper ⢠2609.22586 ⢠Published 23 days ago
RL-ADA: A World-Feedback Framework for Adversarially Robust Enterprise Dialogue Agents Paper ⢠2609.02902 ⢠Published Jul 6 ⢠1
TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models Paper ⢠2607.28896 ⢠Published Jul 30
DuplexWorld: Can voice agents help you get through the day? Paper ⢠2608.10716 ⢠Published Aug 11
World Feedback for Clinical Agents: Diagnosing RL in FHIR Environments Paper ⢠2607.01470 ⢠Published Jul 1
Beyond Next-Token Prediction: An RLVR Proof of Concept for Tool-Use Agents on Atlassian Workflows Paper ⢠2607.01465 ⢠Published Jul 1
MedMosaic: A Challenging Large Scale Benchmark of Diverse Medical Audio Paper ⢠2605.00969 ⢠Published May 28
GAZE:Governance-Aware pre-annotation for Zero-shot World Model Environments Paper ⢠2510.14992 ⢠Published Oct 7, 2025
An Evaluation Study of Hybrid Methods for Multilingual PII Detection Paper ⢠2510.07551 ⢠Published Oct 8, 2025
Scalable multilingual PII annotation for responsible AI in LLMs Paper ⢠2510.06250 ⢠Published Oct 9, 2025
Human + AI for Accelerating Ad Localization Evaluation Paper ⢠2509.12543 ⢠Published Oct 6, 2025
Running Agents Personal Assistant Benchmark š§ Scores a personal assistant by what it did on the device