Research
Dynamo AI’s mission is to accelerate Gen AI adoption in enterprises. By pioneering cutting-edge research, we strive to create robust solutions that empower businesses to leverage next-generation AI capabilities.
Research Articles
Shallow Robustness, Deep Vulnerabilities: Multi-Turn Evaluation of Medical LLMs
MedQA-Followup introduces a framework to test LLM robustness in multi-turn medical consultations, revealing drastic accuracy declines and vulnerabilities overlooked by single-turn evaluations.
GuardFormer: Guardrail Instruction Pretraining for Efficient SafeGuarding
GuardFormer, a compact guardrail model pretrained on synthetic policy data, outperforms GPT-4 and Aegis-LlamaGuard in safety classification with far lower cost.
PrimeGuard: Safe and Helpful LLMs through Tuning-Free Routing
PrimeGuard is a new inference-time guardrail method that boosts both safety and helpfulness by routing queries to tailored model instructions.
Know Thy Judge: On the Robustness Meta-Evaluation of LLM Safety Judges
LLM safety judges are fragile—small prompt shifts or adversarial outputs can drastically misclassify harm, exposing major reliability and robustness gaps.
.webp)