Research

Dynamo AI’s mission is to accelerate Gen AI adoption in enterprises. By pioneering cutting-edge research, we strive to create robust solutions that empower businesses to leverage next-generation AI capabilities.

Research Articles

Shallow Robustness, Deep Vulnerabilities: Multi-Turn Evaluation of Medical LLMs

MedQA-Followup introduces a framework to test LLM robustness in multi-turn medical consultations, revealing drastic accuracy declines and vulnerabilities overlooked by single-turn evaluations.

GuardFormer: Guardrail Instruction Pretraining for Efficient SafeGuarding

GuardFormer, a compact guardrail model pretrained on synthetic policy data, outperforms GPT-4 and Aegis-LlamaGuard in safety classification with far lower cost.

PrimeGuard: Safe and Helpful LLMs through Tuning-Free Routing

PrimeGuard is a new inference-time guardrail method that boosts both safety and helpfulness by routing queries to tailored model instructions.

Know Thy Judge: On the Robustness Meta-Evaluation of LLM Safety Judges

LLM safety judges are fragile—small prompt shifts or adversarial outputs can drastically misclassify harm, exposing major reliability and robustness gaps.

.webp)