📐 Large variety of ready-to-use LLM eval metrics (all with explanations) powered by ANY LLM of your choice, statistical methods, or NLP models that run locally on your machine covering all use cases: Custom, All-Purpose Metrics: G-Eval — a research-backed LLM-as-a-judge metric for evaluating on any custom criteria with human-like accuracy DAG — DeepEval's graph-based deterministic LLM-as-a-judge

